Skip to main content
EMPATIX.
← BACK TO SITE

← All news

Category: Artificial Intelligence & Software Architecture

Beyond chatbots: multi-agent architectures in enterprise software

Bringing AI into business processes means turning a probabilistic model into a reliable component: separate roles, validated outputs, reversible actions and everything traced.

Published on · 4 min read

Topics: #Agentic AI #LLM Engineering #Enterprise Systems #API Integration

Geometric nodes connected by glowing paths that pass through a glass control plane before reaching a solid block

For years, "integrating AI" into a business application meant adding a chat window in a corner of the interface. Useful for answering the odd question, not much use for real work. Today companies ask us a different question: can we entrust a language model with a process, not just a conversation?

The answer is yes, on one condition: stop treating the model as an oracle and start treating it as a system component. A component with an interface contract, permissions, tests and logs.

The problem: a model is not deterministic

By its nature, an LLM is probabilistic. The same request can produce slightly different answers, and every now and then a wrong answer delivered with great confidence. In a chat that is an annoyance; in a workflow that records an invoice or updates an order, it is a risk.

So the goal is not to make the model deterministic, which is not possible, but to build an architecture around it that makes the system controllable: every output is validated, every action is limited and reversible, every step is logged.

Separate roles, not a single do-it-all agent

The first shift in perspective is to drop the idea of a single agent that does everything. In architectures that work in production, the tasks are divided:

  • a planning agent breaks the request down into steps;
  • one or more execution agents call APIs, databases and services;
  • a verifier checks that the result complies with business rules and compliance constraints before it becomes final.

Separating roles has a practical advantage: each agent receives only the context and tools it needs, and each step can be tested and observed on its own.

Diagram: the request flows through planner, executors and verifier; if approved it reaches the system, if rejected it goes to compensation, if critical to human approval

Typed, validated outputs

The model must not return free text to be interpreted, but structured data that conforms to a schema. Libraries such as Pydantic in Python or Zod in TypeScript define the contract: required fields, types, formats, allowed values. If the output does not match the schema it is discarded or regenerated, and it never reaches the rest of the system.

It is the same discipline we apply to any external API: content is not trusted until it has been validated.

Reversible actions, not "magic transactions"

When an agent acts on real systems, such as writing to a database, calling a microservice or sending data to a supplier, you need to ask what happens if a step fails halfway through.

There are no ACID guarantees across different services. The solution is the one known from distributed systems: compensating transactions (the saga pattern), where every action has an inverse action to run in case of error, and idempotent operations, which produce no double effects when repeated. For irreversible or high-impact actions, such as payments and communications to customers, the rule is simple: human approval is required.

Sandboxing and least privilege

Prompt injection, meaning malicious instructions hidden in a document or web page the model reads, cannot be fully eliminated today. Its impact, however, can be limited:

  • code and queries generated by the model run in isolated environments;
  • each agent has credentials with only the permissions it needs, often read-only;
  • external content is treated as data, never as instructions;
  • access to tools goes through controlled interfaces. Standards such as the Model Context Protocol (MCP) help define explicitly what an agent can do.

The principle is the same as in traditional security: if something goes wrong, the possible damage must be small by design.

Observability: always knowing what happened

A system that makes decisions must be able to explain them. For every run we record inputs, intermediate steps and agent decisions, tool calls, validation results, latency and cost. OpenTelemetry, with its conventions for generative AI systems, lets this data flow into the same monitoring tools used for the rest of the infrastructure.

On top of this come automated evaluations (evals): a set of real cases on which the system is measured before every change to the prompt, the model or the tools. Without evals, every update is a leap in the dark.

Where they make sense

Multi-agent architectures work best on processes that are repetitive but not trivial, where content needs to be "understood" before acting:

  • reconciling documents, orders and accounting entries;
  • extracting data from PDFs, emails and legacy formats;
  • routing and first-pass handling of requests and cases.

They are not the right solution when a process is already perfectly described by fixed rules: there, a traditional script is cheaper, faster and more reliable.

How we work

When we design an agent-based system we start from the process, not the model. We map the steps, identify the ones that genuinely require language understanding, and define the data schemas and the human checkpoints. Only then do we choose the model and tools, and build the evaluation set that tells us whether the system is ready for production.

Want to find out whether one of your company's processes can be safely automated with AI? Let's talk.