Building and Evaluating Agent Harnesses · Part 1

Designing an agent harness: from task to architecture

An agent harness is the application code that turns a model’s decisions into work: it supplies context, runs permitted actions, maintains state and checks when the task is complete. If you are building an agent for your own application, designing that code means deciding how much freedom the task needs and what the application must guarantee regardless of the model’s answer.

Those decisions come before the framework. A support assistant, a coding agent and a document classifier need different ways to act, recover and demonstrate success. Giving all three the same loop, memory store and collection of tools would hide the requirements that make them different.

This article follows those decisions from a task description to a first implementation. It extends Harness Engineering for AI Agents, which introduces the responsibilities of a harness. Here we will work through how to choose them for a particular project.

Why the model needs a harness

A model can propose a refund. Something else must identify the account, load the order, decide whether the proposed operation is permitted, call the payment service and record what happened. If that service times out, the application must also decide whether trying again is safe. These responsibilities exist even when the model is very good at choosing the next step.

LangChain uses a broad definition of a harness that includes the code, configuration and execution logic around the model. In practice, some of that may already belong to your application: authentication, database access, a job queue or an approval system. Designing a harness includes deciding how the agent uses those facilities. It does not require rebuilding them inside an agent framework.

The smallest version might make one model call, validate its output and return a result. A larger one might let the model inspect files, execute programs and revise its work across many turns. Both need an answer to the same question: which parts of completing this task can we delegate, and how will we know they happened correctly?

To make that concrete, consider a customer-support assistant that handles refunds, delivery questions and account changes. We will use it as a design example. Its requirements will lead us toward a particular architecture; a different task should lead to a different set of choices.

Describe the result before the agent

“Handle refund requests” leaves most of the engineering unspecified. Suppose a customer asks for a refund on a delivered order. A successful result requires a refund for the correct order and amount, an explanation to the customer, and no unrelated account changes. If the request falls outside policy, a useful refusal may be the correct result. If the order cannot be identified, the assistant should ask for the missing information.

These are different outcomes. Treating every conversation that ends without an error as success would collapse them into one misleading signal.

We also need to know who is asking. An order identifier in a message does not establish ownership of that order. The authenticated account must come from the application, and the service performing the refund must check that the operation belongs to that account. The model can help interpret the request; it cannot establish the caller’s authority by producing a plausible identifier.

This gives us the beginning of a task contract: the changes that count as success, the changes that are forbidden, and the situations that require a question or a handoff. It also exposes requirements that a prompt cannot satisfy by itself. A maximum refund amount needs an enforcing service. A response-time target needs a deadline. A requirement to resume tomorrow needs durable state.

At this point we have enough to ask whether an agent loop is useful at all.

How much of the procedure do we already know?

If every request follows the same sequence—extract an order number, fetch the order, calculate eligibility and explain the result—we can write that sequence directly. The model may handle the language at the beginning and end. It does not need another call to decide whether to fetch an order that the procedure always requires.

An explicit workflow becomes useful when the procedure has required branches or waiting states. A refund might need approval above a threshold. An address change might be allowed only before dispatch. We can represent these stages as a state machine or graph, with application code deciding which transitions are allowed.

An agent loop gives the model more discretion. In a mixed support conversation, the customer may begin with a delivery complaint, reveal that the wrong item arrived, and then request an exchange. The next useful lookup depends on what the assistant discovers. A loop lets it choose an action, inspect the result and choose again. That flexibility comes with more possible paths to inspect, including unnecessary calls, repeated lookups and premature completion.

LangGraph’s workflow and agent guide makes this distinction in terms of predetermined execution paths versus dynamically chosen steps. A graph can express either. “Graph-based” and “agentic” are not mutually exclusive architectures.

The same refund request drawn three ways. A fixed pipeline runs read request, fetch order, check policy and write reply in a set order. An explicit workflow adds a branch that code decides: a refund over the limit waits for approval. An agent loop lets the model choose each tool call from the last result and decide when to reply.The same refund request drawn three ways. A fixed pipeline runs read request, fetch order, check policy and write reply in a set order. An explicit workflow adds a branch that code decides: a refund over the limit waits for approval. An agent loop lets the model choose each tool call from the last result and decide when to reply.

A useful compromise for our support assistant is to let an agent gather information and prepare a proposed change, then pass that proposal to a fixed validation and execution step. The model can explore within the conversation while ordinary code owns the conditions for a write. If we already know this separation is required, we can build the outer workflow first.

Alternatively, we can begin with a bounded agent loop whose tools enforce those conditions. That keeps the orchestration small while we learn which paths the task actually needs. The choice depends on whether explicit stages, approval waits and recovery rules are already requirements.

Inside an agent loop, the model calls read-only lookup tools and finishes with a proposed refund, which the application saves. A fixed code step checks owner, amount and policy. If the checks pass, code issues the refund, the only step with write access. If they fail, the assistant asks the customer or hands off.Inside an agent loop, the model calls read-only lookup tools and finishes with a proposed refund, which the application saves. A fixed code step checks owner, amount and policy. If the checks pass, code issues the refund, the only step with write access. If they fail, the assistant asks the customer or hands off.

There is no need to introduce multiple agents just to draw more boxes. Separate agents become useful when subtasks need different context or permissions, or can proceed independently. They also need a way to combine results and resolve conflicting actions. Two workers that can both refund the same order create a coordination problem that one worker did not have.

Left: a delivery agent and a billing agent can both call the refund API, so one order is refunded twice. Right: both agents are read-only and return findings; a code step combines them and makes one refund call.Left: a delivery agent and a billing agent can both call the refund API, so one order is refunded twice. Right: both agents are read-only and return findings; a code step combines them and makes one refund call.

What language should the agent use to act?

After choosing who controls the next step, we still have to choose how an action is expressed. These are independent decisions: a workflow can contain a code-writing agent, and an open-ended loop can use only a few narrow tools.

For the support assistant, a named operation such as issue_refund(order_id, amount_cents) is a natural interface. Its schema makes the requested action explicit. The tool can return a refund identifier on success or a structured explanation of why the request was rejected. We can inspect and test this contract without asking the model to generate the implementation of a refund.

Some tasks need no tools. A classifier with all relevant input in its prompt can return a validated label. Retrieval becomes useful when the task needs information outside that input. Writes become useful only when completing the task requires changing something. Each capability should have a reason to exist.

Generated code offers a different action language. Imagine asking an assistant to inspect a table, group rows by customer and calculate several totals. Python can express that work in one program, keeping intermediate values inside the execution environment. Expressing every small operation as a separate tool call could require more exchanges with the model.

Hugging Face’s smolagents illustrates both approaches. ToolCallingAgent produces structured tool calls. CodeAgent produces Python that can call the supplied tools, store values and compose operations. A code agent still uses tools; it expresses their composition as a program.

That expressiveness changes what we must operate. We now need to handle syntax errors, execution limits, dependencies and access to files or networks. Whether the reduction in model exchanges pays for those costs depends on the task and model. It is a comparison to run, not an automatic advantage of code.

With tool calls, the model makes three calls and every intermediate result returns into its context. With generated code, the model writes one program that runs in a sandbox, calls the same operations and returns only the totals.With tool calls, the model makes three calls and every intermediate result returns into its context. With generated code, the model writes one program that runs in a sandbox, calls the same operations and returns only the totals.

Several platforms support a hybrid: narrow business tools plus a sandboxed code-execution tool. OpenAI’s Responses API, the Claude API, the Gemini API and AWS Bedrock AgentCore all offer code execution as a built-in tool that runs next to your own functions. Anthropic’s programmatic tool calling goes one step further: code in the sandbox can call the tools you allow, and only selected results return to the model. Anthropic’s guidance keeps direct tool calls as the default and uses the code path for large results and chains of dependent calls. Code-only designs exist too, such as smolagents’ CodeAgent and Cloudflare’s Code Mode.

For our refund assistant, I would begin with typed business operations. If a later task needs substantial calculation, we can add constrained computation without giving that code direct authority over customer accounts.

MCP answers a further, separate question: how should the application connect to tools? A local function is sufficient when one application owns the implementation. MCP can help when tools are served separately or shared across clients. It supplies an integration protocol; validation and access control remain responsibilities of the participating application and server, as the MCP tools specification describes.

Put restrictions where actions happen

A tool description tells the model when an operation is useful. The implementation must decide whether this particular call may run. For a refund, that means checking the authenticated account, the order, the amount, applicable policy and any required approval at the service that performs the write.

Some checks need interpretation

Suppose the customer asks, “Can I return this order?” and the agent proposes a refund. The order belongs to the customer, the amount is valid and the return window is open. Those checks can all pass while the intended action remains unclear: the customer may be asking about eligibility rather than requesting a refund now. We need to interpret the message before deciding what to do next.

This is a useful place to separate three responsibilities. Ordinary code checks ownership, amounts and dates. The main language model conducts the conversation and explains the result. A separate decision model can assess a narrow semantic question, such as whether the customer actually requested the proposed change. This separation gives us a decision we can test independently of the rest of the conversation.

Jev, a decision model from TypeSafe, is one candidate for that role. It takes supplied state and typed questions, and returns structured answers. For this example, its Noul question type returns a probability for a yes/no proposition. I would ask separately whether the customer explicitly requested the refund and whether a later message withdrew that request. TypeSafe recommends making each question specific and combining the answers in code; several independent questions can share one API call. Keeping the two judgments separate makes a mistaken interpretation easier to locate.

The application then combines those results with its mandatory checks. Uncertain or conflicting answers can lead to clarification or human review; a classifier timeout needs an explicit error path too. A positive semantic judgment cannot override an ownership check or substitute for a required confirmation. Jev’s probability is a model output, not an authorization record or an independent check of correctness.

A proposed refund goes to mandatory code checks and to a decision model that asks whether the customer requested it and whether a later message withdrew it. Code combines both: a failed check leads to a refusal, an unclear answer or an error leads to a question or review, and only all checks passing leads to the refund.A proposed refund goes to mandatory code checks and to a decision model that asks whether the customer requested it and whether a later message withdrew it. Code combines both: a failed check leads to a refusal, an unclear answer or an error leads to a question or review, and only all checks passing leads to the refund.

For such a check, I would supply the proposed action, relevant customer messages in order and any tool results needed for the decision, keeping their sources distinguishable. Passing only the latest message could lose a cancellation; passing the entire account history could bury the relevant request. The question, context and fallback behavior are all parts of the design. A small general-purpose model with structured output is another candidate. Their usefulness and cost need comparison on the same decisions.

Restrict access independently of model judgments

The source of an instruction matters. A customer message or retrieved document can contain instructions. That text must not be able to grant new permissions or replace the application’s authoritative policy. If a document tells the agent to export all customer records, the system should lack permission to do so regardless of how convincingly the document phrases the request.

This is also where the sandbox decision belongs. If the model can only call reviewed functions with narrow permissions, a general-purpose code sandbox may add little. The functions still need authorization, validation, timeouts and safe database operations.

Once we allow generated Python, shell commands or executable file edits, we need to decide what that execution can reach before enabling it. Which directories are writable? Is network access required? Which credentials are available? How much CPU time, memory and output can a run consume? A separate process by itself does not answer those questions.

Hugging Face’s secure execution guide describes restricted local execution and more isolated alternatives. Their restrictions differ. A container with broad host mounts and powerful credentials can still expose exactly the resources we intended to protect. Test an attempted forbidden read, write and network request alongside a successful program.

We can postpone a code runner while our task does not need one. Its isolation needs to be ready when we grant code execution. And isolation cannot prevent misuse of an API that we deliberately make available: the refund service must still enforce its own rules.

What survives when a run stops halfway through?

Now suppose the assistant submits a refund, the service creates it, and the response is lost. The model sees a timeout. Repeating the call may create another refund; declaring failure may tell the customer something that is no longer true.

This problem connects state, retries and recovery. We need an application-owned operation identifier, persisted before dispatch and reused when recovering the same intended action. The service must recognize that identifier and return the existing result rather than perform the action again. AWS’s guidance on idempotent APIs explains this pattern. If the service offers no such guarantee, recovery needs a way to check what happened or ask for a human decision before another attempt.

The harness saves operation ID op-7f3, sends the refund, and the service creates refund r-1. The response is lost and the harness sees a timeout. It retries with the same ID, the service returns the existing refund r-1, and the customer is told that one refund exists.The harness saves operation ID op-7f3, sends the refund, and the service creates refund r-1. The response is lost and the harness sees a timeout. It retries with the same ID, the service returns the existing refund r-1, and the customer is told that one refund exists.

Saving the conversation is not enough. It may record that the model requested a refund while saying nothing about whether the external service completed it. A checkpoint helps resume execution; it does not roll back external effects or make repeating them safe.

For this assistant, three kinds of state have different jobs. The conversation holds the customer’s request and the facts gathered so far. The order and refund records describe the business state. A durable execution record tracks pending operations and approvals so a restarted worker can continue safely. They can share infrastructure, but they should not substitute for one another.

Cross-session memory is another choice. If the assistant must remember a preference next week, we need rules for whose preference it is, who can update it and how it can be deleted. If every task is independent, a memory service may be unnecessary. More retained text also means more opportunities to reuse stale or irrelevant information.

The same reasoning applies to summaries. When a conversation outgrows its useful context, we may need retrieval or summarization. A summary that drops an order identifier, approval condition or unresolved operation can break the procedure. Keep essential operational state in explicit fields, and test what survives compression before relying on it.

Give the loop a way to stop

A flexible loop can keep finding another thing to inspect. We therefore need a stopping policy that covers success, missing information, exhausted resources and failure. Reaching a limit should produce a truthful status: the refund may be complete, unattempted or still uncertain. “The agent stopped” is insufficient information for the customer or the next worker.

A model-call limit bounds only part of execution. A tool can hang. A retry can repeat work. A summarizer or guard can make additional model calls. The application needs limits on elapsed time, attempts, output size and spending that include the components it actually invokes. Cancellation also needs to account for work already sent to an external service.

Model choice belongs in this design because it affects what the loop can do within those limits. Test the required action format on representative tasks. Consider usable context, latency, cost and where task data may be sent. A locally served model and a hosted API can implement the same interface while imposing different capacity and operational constraints.

Choose an initial model and endpoint, record their settings, then hold them fixed while evaluating a harness change. Otherwise, a faster run may come from a different provider or model rather than the change we intended to test.

Evaluate completion, including the cases that should do nothing

We can now execute a request, but we still need to decide whether the result is useful. For the refund example, there are at least three observations: what the assistant said, which actions it attempted, and what changed in the service. Each catches failures the others can miss.

“I refunded your order” can accompany an unchanged database. The correct refund can accompany an unrelated address change. An unchanged database can mean a correct refusal, a useful clarification question, a crash or an assistant that did nothing. Final state alone cannot distinguish those last four cases.

Four runs leave the database unchanged: a valid refusal, a clarification question, a crash and a run that did nothing but claims the refund is done. Only the response and the actions tell them apart.Four runs leave the database unchanged: a valid refusal, a clarification question, a crash and a run that did nothing but claims the refund is done. Only the response and the actions tell them apart.

The support simulation accompanying this article makes that limitation concrete. Its twelve custom tasks include seven that expect no database change. Passing an unchanged starting database to each checker therefore passes 7 of 12 state checks. This is a property of the checks, not a measured support success rate. The task definitions inspect final database differences; they do not grade the explanation or intermediate actions. A write that is later reversed can also disappear from that comparison.

For your first evaluation set, include a successful action, a valid refusal, a request missing essential information and a failure during execution. Define the expected response and permitted actions as well as the final state. A clarification case should require the right question; a refusal case should require a useful explanation. These cases expose whether the agent understands the task and whether the application enforces it.

The Jev intent check needs its own cases too. Give it proposed actions with independently assigned expected decisions: an explicit refund request, an eligibility question, a withdrawn request and an instruction embedded in a tool result. That last case also tests whether context assembly preserves the distinction between customer intent and tool-provided text. Count inappropriate actions it allows and legitimate actions it blocks, then follow the full conversation after a block. Preventing a write helps only part of the task if the assistant subsequently loops or never asks the customer to clarify. Choose thresholds on development examples, freeze them before the held-out comparison, and record the question version, supplied context, returned probabilities, applied rule, errors and cost. The recorded demo below did not exercise a Jev rejection, so it cannot establish that benefit.

Record enough to diagnose a failure: the task, model and harness versions, proposed actions, execution results, final state, errors, latency and cost. Keep secrets and unnecessary personal data out of that record. Give the candidate its task input and the rules it needs, while keeping reference answers and held-out test cases inaccessible. The evaluator and its evidence must also remain outside the candidate’s control. If a future optimizer can edit the harness, it must not be able to rewrite the checks or results that judge its changes.

This evidence gives optional components a purpose. Add a router when traces show that task selection needs a separate stage. Try a summary when long histories cause a specific failure. Test a guard against forbidden actions, including cases it should allow. Required access controls follow from the task’s constraints; optional additions need evidence that their benefit justifies their cost.

Turn the design into a first implementation

For our support example, we now have a reasoned starting point: a bounded loop for interpreting varied requests, narrow tools for account operations, authoritative checks in the services performing writes, and explicit records for pending actions. We have no requirement for generated code or persistent conversational memory yet.

I would build one complete path first. Create a small test account and an expected refund outcome. Implement the lookup and refund operations, including their rejection paths. Connect the model to those operations. Then exercise the refusal, missing-information and interrupted-write cases. This exposes missing requirements before a larger task collection makes them harder to see.

A framework-free loop can be sufficient at this stage. The basic flow is to assemble context, request the next action, validate and execute it, then return the result to the model. Owning that loop also means owning message conversion, cancellation, state and tracing. Reusing an existing framework is useful when it saves that work without hiding the decisions we need to control.

For this series, LangChain’s create_agent is a convenient starting point because it provides the model/tool loop and hooks for changing its behavior. It already runs on LangGraph. Moving to an explicit LangGraph workflow later means taking control of more stages around that agent; it does not mean adding a graph to a system that previously had none.

The first design need not contain every capability discussed above. It needs enough to complete its task within its constraints, plus evidence that reveals where it fails. Before writing the implementation, I would put the decisions on one page:

QuestionDecision to write down
What task are we completing?The caller, expected result, valid refusal or handoff, and forbidden effects
Who chooses the next step?Known stages in code; decisions delegated to the model; reason for each
Which checks require interpretation?The semantic question, its context, candidate model and behavior when uncertain
What can the model ask to do?No tools, typed operations, generated code or a combination
Who authorizes and executes it?Allowed resources, enforcing service, approval rules and any sandbox restrictions
What must survive interruption?Conversation state, business records, pending operations and recovery behavior
What limits apply?Model/provider constraints, data handling, calls, time, spending and cancellation
What proves completion?Response, action and state checks; who owns their evidence
What have we left out?Optional components and the failure or new requirement that would justify each

A useful answer is specific enough to change the implementation. “Needs memory” is vague. “Must resume an approved refund after a worker restart without creating a second refund” tells us what to persist and what failure to test.

This is Part 1 of Building and Evaluating Agent Harnesses. Part 2 takes the control-flow question further: when does an explicit workflow improve the handling of a task enough to justify its extra stages? We will build the smallest motivated workflow around the agent and compare it with the loop. Routers, parallel workers and summaries are candidates to investigate, not a predetermined destination. Later parts separate the evaluator from the candidate and measure variation across repeated runs.

Optional: inspect the implementation and its limits

The companion demo is a small, runnable version of the support assistant. It runs in two environments:

  • Custom support tasks: twelve tasks, each against a fresh SQLite database of customers and orders.
  • τ³-bench retail: eight tasks from a public benchmark in which a simulated customer talks to the agent and database checks score the result.

This section describes what the demo contains, what it leaves out on purpose, what the results show and how to run it. You can skip it and still use the design process above.

What the demo contains

One LangChain agent chooses the calls. It has five account tools (look_up_account, list_orders, issue_refund, update_address and set_preference) and an MCP server for policy and FAQ lookups. An in-memory checkpointer holds execution state. There is no code runner and no memory across sessions.

The demo records two configurations:

SpecComponents
plainSummarization, a limit of 20 model calls per run, and 2 retries on a tool error
baseEverything in plain, plus Schema-Guided Reasoning and a Jev write guard

Schema-Guided Reasoning (SGR) makes the model return a structured next-step object instead of a native tool call. The Jev guard runs before each write and asks one compound question: would this call break the policy or go beyond what the customer asked for? The two separate questions earlier in this article, about the request and its withdrawal, are a proposed redesign. The recorded runs did not use them.

base adds SGR and Jev together, so the custom results cannot show which of the two caused a difference.

The model node carries summarization, the model call limit and Schema-Guided Reasoning. The tools node carries tool retry and the Jev write guard. Under τ³-bench the graph stops before the tools node, so only the model-side components run.The model node carries summarization, the model call limit and Schema-Guided Reasoning. The tools node carries tool retry and the Jev write guard. Under τ³-bench the graph stops before the tools node, so only the model-side components run.

What the demo leaves out on purpose

  • The tools check data integrity but leave business policy to the agent, so an experiment can expose a write the policy forbids. A real service would enforce the policy at the write.
  • There is no task-wide limit on time or spending, and retried writes have no idempotency key.
  • The evaluator runs in the same process as the agent.

These results therefore do not test production enforcement, retry safety or protection of the evidence against tampering.

Results on the custom tasks

Each row is one run per task, with the model and endpoint prices recorded with the results. A state check compares the final database with the expected one. Cost and latency are means per task; cost includes Jev where it runs.

SpecModelState checksCostLatency
baseopenai/gpt-6-luna12/12$0.0007311.9 s
plainopenai/gpt-6-luna12/12$0.000335.8 s
basez-ai/glm-5.3-flash12/12$0.001019.7 s
basexiaomi/mimo-v2.6-flash7/12$0.0008718.6 s
  • With gpt-6-luna, both specs passed all twelve tasks. base cost about 2.2 times as much and took about twice as long.
  • Jev allowed every write it checked: twelve across the gpt-6-luna and GLM runs, and four in the MiMo run. It never blocked a write, so these runs do not test whether it stops a bad one.
  • MiMo’s five failures were SGR output-format errors. There is no MiMo plain run, so we cannot say whether native tool calling would do better.
  • Summarization never ran. The largest request in the six recorded runs (these four and the two τ³-bench runs below) had 8,351 input tokens, below the 12,000-token trigger. These results cannot show whether summarization helps.

Results on τ³-bench

τ³-bench scores a run by replaying the agent’s tool calls in a fresh environment, so the benchmark must run the tools itself. The adapter pauses the graph before its tools node, returns the tool calls, and writes the benchmark’s results back before it resumes.

One turn of the τ³ adapter. The model node runs. If the graph is paused before the tools node, the adapter returns the tool calls, τ³-bench runs them, and the adapter writes the results into the checkpoint and resumes. Otherwise it returns the text reply.One turn of the τ³ adapter. The model node runs. If the graph is paused before the tools node, the adapter returns the tool calls, τ³-bench runs them, and the adapter writes the results into the checkpoint and resumes. Otherwise it returns the text reply.

This has two consequences. First, the prompt, tools and policy come from τ³-bench, so these results are not comparable with the custom ones, even with the same spec file. Second, tool retry and Jev never run here: the only difference between base and plain is SGR.

Both runs use openai/gpt-6-luna for the agent and for the simulated customer. Simulation latency includes the simulator’s work. Agent cost and simulation latency are means per task; simulator cost is the total for all eight tasks.

SpecPassedAgent costSimulation latencySimulator cost, total
base6/8$0.0023859.0 s$0.00453
plain7/8$0.0013638.9 s$0.00471
  • The eight tasks come from the retail test split and have no natural-language assertions, so database checks alone decide the reward.
  • Both specs finished all eight tasks without adapter errors. Every failure was a missing write. base failed tasks 9 and 26. plain handed task 27 to a human instead of completing the exchange.
  • A second base run also passed 6/8 but failed tasks 9 and 17. The failing tasks change between runs, so a one-task difference is no reason to prefer plain. These runs do not measure that variation.

Check the recorded numbers

The saved reports and traces keep the per-task outcomes, token counts, configurations, package versions and recorded prices. For the 393 agent chat calls in the six recorded runs, cost calculated from tokens matched the billed amount in each response. This check excludes Jev and simulated-customer calls. The prices are historical, not a quote for your run.

Run it yourself

Install uv and use the tagged version below. The tests run offline once dependencies are installed. The two environment runs make paid model calls through OpenRouter and need an API key.

git clone https://github.com/slavadubrov/agent-harness-lab-public
cd agent-harness-lab-public
git checkout v0.1.2-a1
cp .env.example .env # set OPENROUTER_API_KEY
make test
make a1-custom SPEC=harness/spec/plain.yaml
make a1-tau3 SPEC=harness/spec/plain.yaml

Each run replaces the results under reports/article-a1/. Use git diff to compare them with the committed baseline. The Makefile installs from the frozen uv.lock.

To try another model or tool set, copy a spec under harness/spec/, change those fields and pass its path through SPEC. Provider and custom-component settings accept free-form keys, so validation cannot catch every nested typo.

You can use the design process in this article without running the demo. The demo makes one set of choices inspectable; your own task contract decides which of them belong in your harness.