Loading...

A Tier-1 BFSI enterprise asked us a deceptively simple question: can AI handle routine compliance lookups without creating regulatory exposure? Their compliance teams were spending hours navigating regulations, amendments, and internal policy interpretations to answer everyday product-team questions defensibly. GenAI Protos designed and shipped an agentic AI Compliance Intelligence Platform that interprets natural-language questions, plans a retrieval strategy across authoritative regulatory sources in real time, and returns grounded, citation-bound answers. The result compressed hours of manual cross-referencing into minutes - with the audit trail second lines of defence and regulators expect.
Thousands of regulations, directives, and amendments to track and interpret - most of which interlock and reference each other.
Legal content spread across primary regulator publications, jurisdictional transpositions, secondary guidance, and internal policy memos.
Compliance teams were spending hours cross-referencing legal texts by hand - five browser tabs open to answer one product-team question.
Missing regulatory updates or misreading clauses can cause penalties and reputational damage. The institution needed AI that could reason across documents and prove its conclusions.
We designed an AI-powered system that crawls, searches, and queries the institution’s authoritative regulatory sources in real time to deliver precise, cited answers in plain language. Six capabilities define what the platform does: Natural Language Q&A - analysts ask compliance questions in plain language and get grounded, cited answers. Live Source Access - real-time crawl and query of authoritative regulator endpoints; no stale data, always current. Agentic Reasoning - multi-agent orchestration with real-time search, planning, and self-verification. Citation Engine - every answer linked to the exact article, paragraph, or annex in the source regulation. Compliance-Aligned - region-appropriate data residency, audit trails, and full alignment with applicable data-protection regulation. Custom Tooling - purpose-built tools that query official sources in real time to fetch and reason over live legal content.
A query travels through six steps, from user input to cited answer:
A compliance analyst submits a question in natural language through the web interface.
A planning agent classifies intent and decomposes the query, deciding which authoritative sources to consult and in what order.
Retrieval agents execute real-time calls against authoritative source endpoints; fetched evidence lands in shared workflow state.
The underlying LLM reasons over the live passages and produces a structured draft answer with each clause tagged to its evidence span.
A self-check verifier agent validates every citation against its source passage and flags any uncertainty. Failed clauses are rewritten or escalated.
The verified answer is returned with direct links to the official articles in the source regulation.
The pipeline runs as a deterministic, step-based workflow rather than a free-form agent loop. Compliance Q&A is a repeatable process, and step-based execution produces auditable checkpoints that open-ended ReAct-style traces cannot.
A handful of engineering decisions defined the system.
We built on Agno after evaluating LangGraph, CrewAI, and AutoGen. All of them can be used to build deterministic, step-based workflows; the differentiator for us was that Agno is lightweight and fast. Its small runtime footprint and low per-step overhead mattered in a workload where every query already pays the latency cost of live retrieval plus a multi-step reasoning loop, and where the platform had to scale across many concurrent compliance officers without ballooning infra spend. Agno also ships with a production runtime out of the box (Agent OS), which kept the team focused on the compliance problem rather than rebuilding scaffolding.
Five Pydantic-typed tools wrap the retrieval surface: a primary retrieval tool against authoritative regulator endpoints; a secondary source tool for supplementary guidance; a section / article lookup tool for resolving inside a regulation; a cross-reference tool that follows citations between documents; and a date / version resolution tool that ensures the agent reasons over the version of the rule in force on the relevant date. Narrow, well-typed tools produced cleaner audit trails than any omnibus search tool we prototyped.
Working state is shared across the agent team during a single workflow run. Cross-session memory persists what analysts have asked before and the regulations they tend to reason over. Sessions checkpoint to a Postgres store so an interrupted query can resume cleanly, and retrieved passages are cached in session state - important for both latency and rate-limit hygiene against live source endpoints.
This is where the project lived or died. A post-execution check on the composer rejects any clause without an attached evidence span - no span, no clause. A confidence-scoring guardrail labels low-confidence answers as needs human review rather than answering with false confidence. When authoritative evidence is missing, a human-in-the-loop step routes the query to a named senior analyst rather than improvising. Input-side guardrails (PII detection, prompt-injection defence) run before the planner ever sees the query.
Every LLM call routes through Portkey, our AI gateway and observability layer, which captures full request/response traces, model routing decisions, retries, and cost telemetry per query. Combined with Agno’s workflow-level traces over tool calls, retrieved passages, and verification outcomes, the result is a per-query record of every reasoning step the system took - usable not just by engineers debugging, but by audit and second-line teams reviewing how an answer was derived.
In a regulated industry, the question is not “can the model answer this?” - it is “can the institution defend the answer?” That reframes the entire technology choice.
Agentic systems win in compliance for four reasons. Multi-hop reasoning is native - the agent plans across documents, follows citations, and resolves amendments where a flat pipeline cannot. Auditability is built in - every agent trace is a contemporaneous record of how an answer was derived, exactly the artifact regulators and internal audit want. Tool-use transparency means the agent’s actions are not a black box: queries, passages, and evidence are all logged and inspectable. And graceful degradation - the agent escalates when uncertain rather than hallucinating with confidence. In a consumer app, hallucination is a UX problem. In compliance, it is a regulatory exposure.
The bet underneath the architecture: in regulated industries, the most defensible AI is the one that can show its work - and an agent’s trace is that work.
Workflows beat free-form agent loops for regulated work. Determinism, replayability, and step-level traces matter more than agent autonomy. We started free-form and moved deliberately to a step-based workflow. The system got more reliable and far easier to audit.
Tool boundaries matter more than tool count. Narrow, typed tools produced cleaner traces and more predictable behaviour than any omnibus search tool we prototyped.
Observability is non-negotiable. Without per-query agent traces, debugging a wrong answer is archaeology. Tracing infrastructure was core build, not a nice-to-have.
Evals belong in the build, not after it. Accuracy and reliability evals from week one - LLM-as-a-judge on accuracy, tool-call verification on reliability - caught regressions long before UAT.
Live retrieval beats pre-indexed corpora for live regulation. An index is stale the moment a new amendment publishes. For compliance, real-time grounded retrieval is the architecturally honest choice.
HITL first, autonomy second. Designing the escalation path early was more valuable than chasing full autonomy.
Four Layer Hallucination Guardrails
What Actually Changed For the Compliance Team
Collapsed from hours of manual cross-referencing to minutes for routine queries.
For first-line interpretation queries fell substantially, freeing senior analysts for genuinely novel work.
Every answer is bound to live source passages, so accuracy is a property the system can demonstrate clause by clause.
Full agent traces are captured per query, turning the AI’s reasoning into reviewable artifacts for the second line of defence.
Authoritative regulator endpoints, custom crawl & query tooling, document parsers, metadata extraction.
Multi-agent orchestration (on Agno): intent classification, query planning, and self-verification agents, coordinated through workflows and teams.
Enterprise LLM (Azure OpenAI / Anthropic Claude / open-source as configured) optimised for legal reasoning.
Source attribution, article-level linking, confidence scoring, and answer versioning.
Smart caching for frequently queried regulations, TTL-based invalidation, graceful fallback when source endpoints are degraded.
Web-based Q&A dashboard with inline citation viewer, search history, saved queries, alerts, and a role-based admin panel.
Region-appropriate data residency, SSO/RBAC, immutable audit logging, encryption at rest and in transit.
If you lead engineering, compliance, or AI strategy at a bank, NBFC, asset manager, or insurer, the highest-leverage place to deploy agentic AI is where audit trails are mandatory and answers must be defensible. GenAI Protos builds production-grade agentic AI for regulated industries - grounded by design, observable end-to-end, and engineered for BFSI-grade auditability. We have shipped it for one of the most established banking groups in our market.
We'd love to hear from you.
Deploy agentic regulatory Q&A that delivers grounded answers, live citations, and traceable workflows for BFSI teams.