Compose capabilities around a requirement
A composable AI application connects capabilities such as retrieval, classification, generation, validation, and execution. The design should explain which requirement each component serves and how its output contributes to the final result.
More components do not automatically produce better quality or efficiency. Each boundary adds contracts, failure behavior, and operational work. Begin with a representative workload and a simple baseline, then add complexity when evidence supports it.
Separate the architectural choices
Modular code, multiple models, independently deployed services, and multiple agents are different choices. Several modules may run in one application. One service may use several models. Multiple services do not necessarily require dynamic agent coordination.
Choose boundaries from ownership, release needs, resource constraints, and behavior. A specialist model may suit a narrow task, but its value depends on measured performance and integration costs. Specialization alone does not prove an advantage.
Understand workflow and service tradeoffs
Anthropic's workflow guidance describes patterns including chaining and routing, with programmatic checks between steps. Use these as design options rather than a requirement to make every application agentic.
The Microsoft microservices architecture guide discusses independent services alongside distributed-system complexity. Review those tradeoffs when deployment boundaries are needed. An AI pipeline can remain modular without adopting a separate service for every stage.
Specify intermediate contracts
Define inputs and outputs, permitted values, missing-data behavior, errors, and version compatibility for every boundary. Include units and provenance where they matter. A valid JSON object is not proof that its contents are correct.
A retrieval stage should preserve source identifiers and access scope. A drafting stage should distinguish supplied evidence from unsupported inference. A validation stage should return actionable reasons for rejection rather than a vague quality score.
Test routing as a decision
A router chooses which path handles an input. Its mistakes can send a difficult case to an unsuitable model or expose a tool outside the intended scope. Define routes, ambiguous cases, and fallback behavior explicitly.
Measure routing errors on representative held-out inputs and inspect consequential misroutes separately. Include unfamiliar requests, missing context, and inputs that should be rejected or reviewed. Compare the router with a simpler rule or single-path baseline.
Keep authorization across boundaries
Pass the identity and access context required for each stage, and enforce permission checks at the point of access or execution. Do not assume an earlier component's success authorizes a later write. Avoid broadening access merely to simplify integration.
Track sensitive information through the pipeline. Decide which components need it, what may be logged, and how retention or deletion applies. Retrieved content and intermediate model output should not override host authorization rules.
Design failure and fallback behavior
Specify what happens when a component times out, returns incomplete output, rejects a request, or produces incompatible data. A fallback must satisfy the required behavior or return a defined failure state; silently replacing evidence with an invented answer is not recovery.
Bound retries and inspect authoritative state after uncertain writes. Test duplicate requests and partial completion. Keep a trace of the stages that actually completed so the application can explain an incomplete result and resume appropriately.
Measure the complete path
Evaluate final task quality alongside component diagnostics. A retrieval improvement can be lost during generation, and an accurate classifier can still create a poor user outcome through a mistaken downstream action. End-to-end tests expose these interactions.
Measure latency, usage, retries, review work, and cost for the actual route distribution. Include network and orchestration overhead. Parallel execution may reduce elapsed time for independent work, but it can increase total resource consumption; measure both.
A hypothetical document-processing pipeline
Suppose a team processes incoming supplier documents. A pipeline retrieves an authorized file, extracts fields, checks required values, and prepares a draft record for review. This is an illustrative architecture, not a reported benchmark or client outcome.
Start with one deployment containing separate modules. Test scanned files, missing fields, conflicting totals, permission failures, and interrupted execution. Preserve supporting text and file identity so a reviewer can check the draft against its source.
Consider a separate extraction service only if measured load, ownership, or release requirements justify it. Compare the proposed change with the existing deployment using the same documents and acceptance criteria. Document added operational costs and unresolved failures.
Version components together where necessary
Record prompts, models, schemas, retrieval settings, and tool contracts in the release configuration. Test compatibility while old and new versions overlap. A component update can change downstream assumptions even when its transport format remains valid.
Define rollback and data-migration behavior together. Rolling back code may not reverse a completed write or schema change. Keep release evidence and a recovery process that the operating team can actually execute.
Make the decision reviewable
A useful architecture record states requirements, candidates, baseline results, tradeoffs, selected boundaries, and review triggers. Explain what additional composition enables and what it costs. Avoid describing the design as universally scalable or future-proof.
Review when observed requirements change: an unmet latency objective, recurring integration failure, independent ownership, or a new access boundary. Popularity alone is weak evidence for replacing a working architecture.
Read our technology selection guide and MLOps guide, explore AI consulting, or contact TensorBlue with the workload and architectural decision you need to resolve.