Inference settings — NuPaaS Docs
AI

Inference settings

The inference gateway serves chat, embeddings and transcription from one endpoint. This page explains the settings that control it, how your own provider keys work, and which numbers are estimates. The request and response formats of the gateway are in the inference API reference.

The gateway and its keys

The gateway lives under /api/inference/v1. It accepts an API key that has the scope inference:run. That scope runs inference and nothing else. A key with only inference:run cannot read or change a setting, a key, a guardrail or a harness.

The settings on this page live under /api/v1/ai. They need a key with the scope read, write or admin. The organization always comes from the key. No request field names an organization.

ParameterTypeDescription
readscopeRead the config, the model list, usage, quota, BYOK rows, guardrails and prompts.
writescopeChange the config, guardrails and prompts. Run a prompt.
adminscopeSave, change or delete a provider key. Delete a harness.
inference:runscopeCall the gateway. No access to the settings.

Settings

A setting is either enforced or not enforced. The gateway reads an enforced setting on each request. A setting that is not enforced has no reader. The platform refuses to save a value for it, so a stored value never looks like a rule that does not exist.

Read the current state from the server. The call returns each setting and an enforcement map with one entry per setting.

ParameterTypeDescription
allow_listenforcedThe models your organization may call. A request for any other model is refused.
default_modelenforcedThe model used when a request names none. The same allow-list and quota rules apply to it.
fallback_modelenforcedThe model tried once when the platform providers fail. The response header x-inference-fallback names the model that was asked for.
prevent_overridesenforcedA chat request that names a model other than the default is refused with 403 model_override_not_allowed. It needs a default model.
inference_enabledenforcedThe switch for the whole gateway for your organization.
output_index_enabledenforcedWhether the platform indexes your served output so that a stripped assertion can still be identified.
retention_daysenforcedHow long the output index keeps an entry. The platform window of 90 days is the upper limit. Usage records are billing records and do not follow this setting.
monthly_budget_centsenforcedThe gateway checks the organization budget before each chat request. The figure is an estimate at the plan overage rates. Null means no budget.
include_byok_spendenforcedWhether tokens that go through your own provider keys count against the organization budget.
cost_tiernot enforcedNot enforced. The platform has no model price tiers. Only the value balanced is accepted.
provider_sortnot enforcedNot enforced. Each model has one platform provider, so there is nothing to order. Only the value balanced is accepted.
capture_content_enablednot enforcedNot enforced. The platform does not capture prompts or responses. Only false is accepted.

Your own provider keys

You can save a key from your own provider account. The vendor bills you for those requests. NuPaaS does not bill the tokens and adds no fee. The gateway shows only the last four characters of a saved key. It never returns the key.

Each key has a priority. The priority decides when the gateway uses it.

ParameterTypeDescription
alwayspriorityThe only route for models the key serves. A failure is returned as it is. The platform providers and the fallback model are never called.
preferpriorityTried first. A rate limit, an outage or a refused key moves the request to the platform providers.
fallbackpriorityTried after the platform providers are used up.

The gateway can route to these providers: OpenAI, Mistral, DeepSeek, Groq, Together, Fireworks, Cerebras, DeepInfra and xAI. It can also route to any OpenAI-compatible HTTPS endpoint that you give as a custom provider. A key for any other provider is refused when you save it. A key counts as routable only when the provider lists the model for your key.

Guardrails and prompts

A guardrail is a policy for API keys: allowed models, a monthly budget, personal-data redaction, retention and moderation. One guardrail can be the default for keys that have none. You can assign a guardrail to one key. Assigning none returns the key to the default policy.

A saved prompt is a template that you can run through the gateway. A run passes the same checks as a chat request and uses tokens. The result of a run is not stored.

Budgets, quotas and rate limits

Three limits can stop a request. Each one has its own refusal.

ParameterTypeDescription
Plan quotahard limitThe token limit of your plan for the billing period. A request over it is refused.
Key budgetestimateA budget on one API key, from its guardrail. The gateway prices the tokens at the plan overage rates and compares them with the budget.
Organization budgetestimateThe monthly_budget_cents setting. It is priced the same way and answers 429 inference_org_budget_exceeded.
Rate limitper organizationA limit on requests per period. Chat requests and prompt runs count against it.

Provenance marks

The gateway can mark generated output with a signed assertion in the field platform_provenance. A prompt run returns the mark unchanged. If a call returns no generated output, the field is null. Verify a mark with the public key and the verify endpoint. The reference explains the claims and the limits of a mark.

Usage

Usage is the input and output tokens of the billing period. Pass a harness run id to see only the rows that the gateway tagged with that run. A harness key sends the run id in a header that the gateway sets. The tag counts only calls that carried it.

From the SDK and the CLI

These calls use the public API, so they follow the scopes above.

SDK
import { PlatformClient } from "@type-driven/platform-sdk";

const client = PlatformClient.fromEnv();

const config = await client.ai.getConfig();
console.log(config.enforcement.cost_tier); // "not_enforced"

await client.ai.updateConfig({ default_model: "mistral-small-latest" });

await client.ai.putByokKey("openai", {
  label: "main",
  api_key: process.env.OPENAI_API_KEY!,
  priority: "prefer",
});

const usage = await client.ai.getUsage({ harness_run_id: "RUN_ID" });
console.log(usage.total_tokens, usage.tagged_rows);
CLI
platform inference config get
platform inference config set --default-model mistral-small-latest
platform inference byok list
platform inference byok put openai --label main --key-file ./openai.key --priority prefer
platform inference byok delete openai --yes
platform inference guardrails list
platform inference prompts run --user "Write one full sentence."
platform inference usage --harness-run RUN_ID

The key file holds the provider key. The CLI never takes a key as a command argument, so the key stays out of your shell history.