Applied AI

Your AI prototype works. Is the workflow ready for production?

A practical way to scope an AI workflow, evaluate real failures, design human review and release a system that can be paused safely.

Dopstack Technologies4 min read
A vivid decision gate sends an uncertain AI recommendation to human review instead of an automated action.
A vivid decision gate sends an uncertain AI recommendation to human review instead of an automated action.

The short version

A production AI workflow needs a narrow permission boundary, a representative evaluation set, an owned review queue and a reversible release. A good demo establishes none of those on its own.

Define one decision and its permission boundary

A model that reads a claim, extracts a policy number, recommends a queue and approves payment is performing four different jobs. The first three might support an intake reviewer; the fourth changes a customer's outcome. Treating them as one 'AI automation' hides which error matters and who can correct it.

Start with the smallest useful action: for example, produce a structured intake draft with a link to each supporting document. Specify permitted inputs, output fields, downstream systems and a hard list of actions the workflow cannot take. A draft-only pilot is less impressive than autonomous approval, but it gives the team a testable boundary and keeps financial decisions with an authorized person.

NIST's AI Risk Management Framework calls for defined tasks, application scope and human oversight. That is a governance principle, not proof that a particular implementation is safe. Your team still needs to set the risk tolerance for its own workflow.

Evidence: NIST: AI Risk Management Framework Core

Decision boundary

Separate the useful draft from the high-stakes action.

01

Claim arrives

Documents and context enter intake.

02

AI drafts

Extract fields and suggest a queue.

03

Human decides

Review, correct and approve any consequential action.

A bounded claims-intake workflow keeps approval with a person.

Evaluate outputs and the path taken to get them

Build a test set from the records the workflow will actually encounter: incomplete scans, contradictory dates, multiple policies, a missing attachment and a request outside the authorized scope. For each case, record the expected fields, evidence and next action, including 'stop and ask a reviewer.' Keep test cases separate from examples used to tune prompts.

Google Cloud distinguishes final-response evaluation from trajectory evaluation, which examines the sequence of tool calls. That distinction matters when an answer looks plausible but came from the wrong document or an unauthorized lookup. For a simpler extraction pipeline with no agent tools, field-level and evidence checks may be enough; do not add agent machinery just to gain another metric.

Measure the cost of mistakes in operational terms: wrong routing, omitted evidence, reviewer correction time and cases that should have stopped. Repeat the evaluation whenever prompts, models, tools or source documents change materially. A single launch-day score cannot describe an evolving workflow.

  • Use held-out cases, including failures and out-of-scope requests.
  • Check source evidence as well as the proposed answer.
  • Track reviewer corrections and incorrect routing.
  • Define what result blocks release.

Evidence: Google Cloud: Evaluate Gen AI agents · NIST: AI Risk Management Framework Core

Evaluation set

Test the records a polished demo hides.

Input conditionExpected action
Missing pagesStop and request evidence
Conflicting datesFlag the conflict for review
Poor scanRoute for manual correction
Each awkward input needs an expected response, including a deliberate stop.

Make human review a real operating path

A label saying 'low confidence' is not a handoff. The reviewer needs the original record, extracted fields, cited evidence, reason for escalation and a way to correct the result without starting over. Assign an owner to the queue, decide how overdue cases surface and make sure the model cannot silently act while a case is pending.

Set routing thresholds using observed performance and the cost of error, not the model's self-reported certainty. Some conditions should be unconditional stops: a missing required document, a policy mismatch or a tool response that cannot be traced. This is an engineering recommendation; the precise stop rules depend on the business and regulatory context.

Evidence: NIST: AI Risk Management Framework Core

Release in stages, with a switch back

First, run the workflow in shadow mode beside the existing process and compare drafts with reviewed outcomes. Then allow staff to use it as a recommendation system while they retain the final action. Only broaden permissions after the team understands observed failure modes, review load and exception handling.

Before any automated write, document how to disable it, where unfinished work goes, how duplicate requests are detected and how a disputed result is reconstructed. Monitor latency, cost, error categories and review backlog together. A faster model that overwhelms the reviewers is not a successful deployment.

The production milestone is not a model endpoint. It is an owned workflow that can be measured, audited, corrected and paused when its assumptions stop holding.

Evidence: NIST: AI Risk Management Framework Core

Sources & further reading

Written by Dopstack Technologies

We design and build software, cloud infrastructure and AI workflows. These notes explain engineering decisions; illustrative scenarios are not claims of client results.

Meet the team

Keep exploring.

Put the idea to work

Start with the problem in front of you.

Tell us where your system gets difficult. We’ll help you map the constraints and a practical first step.

Talk to an engineer