Harness
A harness is an agent that works on a repository in a loop. A maker model writes the change and a checker model reviews it. Each organization can run one harness. It calls models only through the inference gateway of your organization.
What a harness is
A harness runs as a workload in your organization. It has its own key for the gateway, so its calls follow your allow-list, your guardrails and your quota. A run is one pass of the loop over the target repository, with a token budget and an iteration cap.
Deploy a harness
Create the harness in the panel, or with the CLI or the SDK. A create that carries the model map mints the credentials of the harness at once. You can then install it with no extra save.
platform harness deploy --name build-agent \
--repo https://git.example.com/acme/app --ref main \
--maker-model mistral-large-latest --checker-model mistral-small-latest
platform harness install HARNESS_ID
platform harness get HARNESS_IDimport { PlatformClient } from "@type-driven/platform-sdk";
const client = PlatformClient.fromEnv();
const harness = await client.harness.createHarness({
name: "build-agent",
target_repo: "https://git.example.com/acme/app",
target_ref: "main",
model_map_json: JSON.stringify({
maker: "mistral-large-latest",
checker: "mistral-small-latest",
}),
});
await client.harness.triggerHarnessOperation(harness.id, { operation: "install" });
const run = await client.harness.startHarnessRun({ token_budget: 200000 });
console.log(run.loop_id, run.status);The models
The harness needs a maker model and a checker model. The platform checks both when you save the config, before it writes anything:
- The map must be valid. The error names the field that is wrong.
- Your organization policy must allow each model: the allow-list, the inference switch and the default guardrail.
- The gateway catalog must list each model as a chat model it can serve.
If the catalog is not available at that moment, the save is refused and you can retry. A harness without a saved model map cannot install. The install is refused and the answer lists the missing names.
Install, restart and teardown
| Parameter | Type | Description |
|---|---|---|
| install | operation | Starts the harness in your organization. It is refused if the config lacks a value it needs. |
| restart | operation | Rolls the pod and applies the saved config. It also loads a key that the platform minted again. |
| teardown | operation | Deletes the harness. It needs the admin scope and, in the CLI, the flag --yes. |
Each operation answers accepted at once. The work runs on its own. Read the harness again to see the result.
A teardown checks first that the resources belong to this harness. Then it removes the stack and its network policy, revokes the key of the harness, removes the forge user and its access, revokes the Plane token and deletes the Plane project, deletes the stored credentials and deletes the config. It deletes the credentials for good. It keeps no copy. To use the harness again, install a new one. The platform mints new credentials. After a teardown the name is free to use again.
Why a run stopped
A run that stops for a reason carries a halt reason. The harness page shows it in words. These are the reasons the platform knows and the fix for each.
| Parameter | Type | Description |
|---|---|---|
| gateway_quota_exceeded | halt reason | A limit was reached: the harness token budget, the organization quota or the organization budget. Raise one of them or wait for the period to reset. Then start a new run. |
| gateway_rate_limited | halt reason | The gateway rate-limited the harness and the retries ran out. Start a new run in a few minutes. |
| gateway_policy_denied | halt reason | A guardrail or a data marking rule refused a request. Change the guardrail or change what the harness sends. Then start a new run. |
| gateway_model_unavailable | halt reason | The gateway cannot serve the maker or checker model. Pick a model the gateway serves, save the config and start a new run. |
| gateway_auth_failed | halt reason | The gateway refused the key of the harness. The platform mints this key again by itself. Restart the harness to load the new key. Then start a new run. |
| gateway_provider_unavailable | halt reason | Every provider that can serve the model failed. The fault is at the provider. Start a new run when the provider recovers. |
A reason that is not in this list is shown as plain text from your own harness, cut to a safe length.
The three limits
Three limits can stop a harness. A stop for any of them reads gateway_quota_exceeded.
| Parameter | Type | Description |
|---|---|---|
| Harness token budget | per run | The token_budget of the harness or of one run. The harness stops the run when it is used up. |
| Organization quota | per period | The token limit of your plan. It is shared with all other inference in your organization. |
| Organization budget | per period | The monthly_budget_cents setting of the inference config. It is an estimate at the plan overage rates. |
The policy posture
The platform reports a policy posture for the harness: enforced or not enforced, and the reason. A posture reads enforced only when the policy mode is on, the policy gateway is ready in the namespace of the harness, and the harness is serving. The policy mode in the config is your intent. It is not the posture.
What a run costs
The cost of a run is an estimate at the plan overage rates. It is not an invoice. The tokens come from one of two sources, and the answer always says which.
| Parameter | Type | Description |
|---|---|---|
| run_tag | attribution | The usage rows that the gateway tagged with the id of this run. A harness key sends the id, so a call by a person in the same window does not change the figure. A call without the tag is not in the sum. |
| org_window | attribution | No row carries the run id. The figure is the whole inference usage of your organization in the time window of the run. It includes other calls. The answer carries a note that says so. |
platform harness run list
platform harness run cost RUN_ID
platform inference usage --harness-run RUN_ID