TensorBlue Blog
AI & Innovation
AI & Innovation8 min read

AI Agents: Tool Use, Permissions & Evaluation

TensorBlue TeamUpdated 8 min read

Design AI agents around bounded tasks, explicit tool permissions, persisted state, stopping conditions, and evaluation of actions as well as final answers.

Define what the agent may accomplish

An AI agent can select actions and use their results to decide what to do next. Its usefulness depends on the task, tools, execution environment, and evidence that the task was completed. A convincing final answer does not establish that the correct actions occurred.

Start with a bounded objective and observable completion criteria. Identify the permitted resources, the user whose authority applies, and the actions that require review. Define what happens when information is missing or the task cannot be completed within the available budget.

Choose a workflow or an agent deliberately

Anthropic's engineering guide distinguishes predefined workflows from agents that dynamically direct their process and tool use. It also discusses added cost, latency, and compounding errors. The article's conceptual patterns are useful; verify current tooling documentation before choosing a particular implementation.

A fixed workflow may suit a predictable sequence such as extracting fields, validating them, and preparing a draft. Dynamic action selection may help when the next step depends on discoveries during the task. Compare both approaches against the same acceptance criteria rather than assuming greater autonomy is better.

Make tools precise and narrow

Give each tool a clear purpose, argument contract, response format, and failure behavior. Distinguish a read from a write and a proposed action from an executed one. Include enough result information to let the system assess progress without exposing unrelated data.

For example, an order lookup should return an authorized order's relevant state and freshness. A refund proposal should describe the requested change. A refund execution tool should enforce the application's policy independently. Do not hide several consequential actions behind a vaguely named utility.

Enforce authorization in the application

Validate the authenticated user's rights, target resource, permitted destination, and action scope before execution. Recheck authority when an action is performed. A model-generated argument or instruction is not proof of permission.

Treat documents, retrieved pages, and tool responses as untrusted input. They may contain instructions that conflict with the user's task. Keep credentials outside model context and limit the data available to the task. Prompt instructions can help express boundaries, but application controls must enforce them.

Persist state without claiming automatic learning

Stored conversation history, retrieved records, and task state can provide context across steps. This is different from updating a model's weights. An agent that remembers an earlier result is not necessarily learning a new general capability.

Define which state is retained, who can access it, its expiry, and how corrections or deletion are handled. Track the origin and freshness of remembered facts. Avoid carrying a previous user's permissions or private records into a new task.

Handle retries and partial completion

A timeout does not prove that an external write failed. Before retrying a consequential action, inspect its authoritative state where possible. Use appropriate request identifiers and duplicate handling in the host application so a retry cannot silently repeat a transaction.

Record proposed, approved, attempted, completed, and failed actions separately. If a task stops midway, tell the user what actually changed and what remains. A resumable task should recover from recorded evidence, not infer completion from an earlier plan.

Set stopping conditions and escalation paths

Bound execution by time, tool calls, cost, or task-specific constraints. Stop when completion criteria are met, further attempts cannot help, or the defined review condition is reached. Detect repeated actions and unchanged results rather than allowing an indefinite loop.

Make escalation useful: provide the relevant evidence, uncertainty, proposed next action, and responsible reviewer. Avoid asking for the same approval repeatedly when it still applies, but reassess when the requested scope or destination changes.

A hypothetical support workflow

Suppose an authorized customer asks why an order has not arrived. An agent retrieves the order state, checks the relevant delivery record, and drafts an explanation supported by those records. This is an illustrative design, not a measured client result.

If the records conflict, the agent routes the case for review rather than inventing a delivery date. If the customer requests a refund, the application checks eligibility and the required approval before any write. A tool response confirms the result, and the final answer describes only what was verified.

Test missing orders, stale delivery records, instructions embedded in a support document, expired authorization, duplicate requests, and interruption after a write. These cases test the task's real boundaries more effectively than a demonstration that assumes every dependency succeeds.

Evaluate the trajectory and the result

Use representative held-out tasks with an explicit rubric. Inspect tool selection, arguments, authorization checks, source use, stopping behavior, and the final artifact. A correct-looking answer can conceal an unauthorized read or an unsuccessful action.

Compare with a simpler baseline. Record task completion, consequential errors, review rates, latency, usage, retries, and cost under comparable conditions. Separate development examples from final evaluation and include failures that matter to the application.

Release with traceability

Version prompts, tools, policies, model configuration, and evaluation results together. Retain a known working release and define rollback conditions. Redact sensitive data from operational logs while preserving the evidence needed to investigate failures.

Monitor actual task outcomes and changing inputs. Offline evaluation supports a bounded release decision; it does not establish permanent reliability. Expand permissions or task coverage only after reviewing the additional requirements and evidence.

Read our prompt engineering guide and MLOps guide, explore MCP server development, or contact TensorBlue with your task, tools, and acceptance criteria.

Tags

AITechnologyInnovation
T

TensorBlue Team