TensorBlue Blog
AI & Innovation
AI & Innovation10 min read

Prompt Engineering: Examples, Validation & Evaluation

TensorBlue TeamUpdated 10 min read

Build useful prompts with clear tasks, bounded inputs, worked examples, output validation, representative evaluations, and controlled production changes.

Define the task before changing the wording

Prompt engineering is the work of specifying a model's task and testing whether the resulting system performs it reliably enough for its intended use. A longer prompt, a named technique, or a confident answer does not establish quality. Begin with the decision or artifact the application needs, the evidence available, and the consequences of an incorrect result.

For an extraction task, define fields and missing-value behavior. For support drafting, define permitted sources and escalation conditions. For classification, define labels and ambiguous cases. Keep the task narrow enough that a reviewer can distinguish an acceptable answer from an attractive but unsupported one.

Separate instructions, context, and input

State the task, constraints, input boundaries, and output contract clearly. Mark source material with explicit delimiters or separate application fields. Explain which material provides evidence and which instructions govern the task. This makes the request easier to inspect and maintain.

The OpenAI prompt engineering guide discusses instructions, examples, context, and message roles. Apply those ideas to the actual API and model being used. Role names and supported parameters can differ across providers and interfaces; verify the current documentation for the deployed configuration.

Do not put credentials into a prompt or rely on a sentence that says an action is forbidden as the application's authorization system. Retrieved pages, uploaded documents, and user input may contain adversarial instructions. Treat their contents as untrusted data and enforce access and action permissions in application code.

Use examples to define important boundaries

A worked example can clarify the format and the distinction between similar labels. Include ordinary cases and relevant edge cases: missing information, conflicting evidence, ambiguous language, or inputs outside the task. Examples should follow the same rules as the instructions.

Choose examples that demonstrate behavior rather than teaching accidental shortcuts. If every urgent example contains a particular customer name, the model may associate that name with urgency. Review examples for irrelevant correlations and confidential information before including them.

Keep evaluation examples separate from prompt examples. Repeatedly editing a prompt until it passes the same small set creates weak evidence of generalization. Maintain a held-out set and record which cases influenced development.

A hypothetical extraction prompt

The following is an original illustration of a delivery-date extraction task. It is not a measured API result or a guarantee that any model will return the expected answer.

Task: Extract an explicitly stated requested delivery date.
Use only the message inside INPUT. Do not infer dates.
Return requested_delivery_date as YYYY-MM-DD, or null if absent.
Return evidence as the exact supporting phrase, or null if absent.
If the message contains conflicting requested dates, return null
for both fields and set needs_review to true.

INPUT:
Please deliver order A17 on 2026-11-04.
END INPUT

The expected output for this example is:

{
  "requested_delivery_date": "2026-11-04",
  "evidence": "on 2026-11-04",
  "needs_review": false
}

Now test a message containing no date, two conflicting dates, an invalid calendar date, a quoted earlier request, and an instruction inside the input asking the model to ignore the task. Decide explicitly whether relative dates are supported and which timezone or reference date would be needed. Do not silently introduce those assumptions.

Validate structure and meaning separately

When a supported model and interface provide schema-constrained output, use a schema that matches the application contract. The OpenAI Structured Outputs documentation describes schema support and handling considerations, including refusals and incomplete responses. Check supported schema features and response states rather than assuming every response is a usable object.

Schema conformity does not establish factual correctness. A well-formed date may still be absent from the input, and a valid category may still be wrong. Validate dates, permitted values, required evidence, and business rules independently. Where practical, compare an extracted supporting phrase with the supplied text.

Handle missing fields, refusal, interruption, and validation failure deliberately. Route uncertain cases to review or return a defined failure state. Do not substitute a fabricated value just to satisfy a required field. Bound retries and capture why a retry occurred.

Build an evaluation set from the real workload

Collect representative, appropriately authorized examples from the task distribution. Include common cases, high-impact mistakes, long inputs, formatting differences, missing information, and adversarial content relevant to the application. Record sampling limitations; a convenient set of clean examples can conceal important failure modes.

Define the rubric before comparing prompts. For extraction, score each field and evidence support. For classification, examine per-label errors and confusion between consequential categories. For drafting, assess factual support, required content, prohibited claims, and whether escalation was appropriate.

The OpenAI evaluation best practices guide emphasizes task-specific evaluation and iterative testing. Use a reproducible evaluation workflow with retained inputs, configurations, outputs, and judgments. The workflow need not depend on a particular hosted evaluation product.

Compare changes with a consistent method

Run the baseline and candidate against the same held-out cases and scoring rules. Record the model identifier, prompt version, relevant settings, retrieval configuration, and tool behavior. Otherwise a change in the surrounding system can be mistaken for a prompt improvement.

Inspect individual failures as well as aggregate scores. A candidate may improve easy cases while worsening the mistakes that matter most. Have qualified reviewers assess consequential judgments and ambiguous ground truth. If automated grading is used, check its agreement with human judgments on an appropriate sample.

Measure latency, input and output usage, retry frequency, and end-to-end cost alongside quality. Use current provider prices when calculating costs and document the assumptions. This guide does not assert universal improvement percentages, fixed token prices, or a guaranteed cost reduction.

Choose generation settings from observed behavior

Verify which parameters the selected model supports before using them. Lower temperature, where supported, can change sampling behavior; it does not guarantee factual accuracy or deterministic output. Test repeat runs when variability matters to the task.

Request the answer and a concise explanation or supporting evidence when the user needs one. An explanation is not proof that the answer is correct. Avoid making private step-by-step reasoning a required production artifact; assess the result against observable evidence and the task rubric.

Check context and output limits for the actual model version. Longer input can add irrelevant or conflicting material, and truncation can remove essential evidence. Test what happens when an input exceeds the application's permitted size and define a visible, controlled failure behavior.

Keep tool execution under application control

A model-proposed tool call is a request for the application to consider. Validate its arguments, the user's authority, allowed destinations, and the action's scope before executing it. Keep secrets in the appropriate credential system rather than exposing them in model context.

For a support workflow, retrieving an order and refunding it are different permissions. Check the specific order belongs to the authorized account and require the configured approval for a refund. Recheck permissions at execution time and make duplicate handling explicit for actions that may be retried.

Test hostile text in documents and tool responses. Evaluate whether it causes data disclosure, unauthorized actions, or incorrect source selection. Prompt wording can contribute to defense, but the application must enforce boundaries independently.

Release, monitor, and revise deliberately

Version prompts with their evaluation results and configuration. Keep a known working version available and define rollback criteria. A release record should explain the intended improvement, observed regressions, unresolved limits, and the owner of the decision.

Monitor validation failures, refusals, review rates, tool errors, latency, and task-specific quality signals. Review authorized samples over time because user inputs and connected data can change. A successful offline evaluation is evidence for a bounded release decision, not a permanent guarantee.

Read our MLOps guide for release and monitoring practices, and our vector database comparison when retrieval is part of the task. Explore LLM inference or AI consulting, or contact TensorBlue with example inputs, the required output, and the errors your application must avoid.

Tags

prompt engineeringGPT-4ClaudeLLM optimizationfew-shot learning
T

TensorBlue Team