Skip to main content
POST
Create an eval run (async)

Authorizations

Authorization
string
header
required

MCPJam API key (sk_…). Create one at Settings → API keys. Guest sessions cannot use the API, and API keys cannot manage other API keys.

Path Parameters

projectId
string
required

ID of the hosted project that contains the server.

Body

application/json

Two valid shapes: suiteId (rerun an existing suite, optionally upserting inline tests into it) or suiteName + tests + serverIds (create a new suite and run it). Inline tests alone — without a suiteId or a suiteName — are rejected with VALIDATION_ERROR.

environmentId requires suiteId: an environment is launchable only through a suite that has it attached (environmentIds, set via PATCH /eval-suites/{suiteId}), so an environment run on a not-yet-created suite could never be satisfied. environmentId and serverIds are mutually exclusive.

suiteId
string
required

Existing suite to rerun. A bare suiteId with no tests reruns the suite exactly as configured.

suiteName
string

Name for a new suite. Required (non-empty) when no suiteId is given.

suiteDescription
string
tests
object[]

Inline test cases to upsert into the suite before running.

Maximum array length: 100
serverIds
string[]

Servers (by ID) the run connects to. Required when creating a new suite; optional on reruns — when omitted, the run connects the suite's saved server selection (the set its snapshot references). A rerun of a suite with no saved selection is rejected with VALIDATION_ERROR (details.reason: "NO_SAVED_SERVER_SELECTION"). Rejected outright for a suite with attached environments (details.reason: "ENVIRONMENT_SERVERS_NOT_OVERRIDABLE"): the environment supplies a closed set that a server override cannot change, so accepting one would connect a different set than the run is stamped with.

Minimum array length: 1
serverNames
string[]

Optional display names, parallel to serverIds.

suiteRerun
boolean

When true, skip per-case upsert and rerun the persisted suite. A bare suiteId with no tests is always treated as a rerun.

modelApiKeys
object

Optional per-provider model API keys (e.g. { "anthropic": "sk-ant-…" }). Falls back to your organization's configured providers when omitted.

notes
string
passCriteria
object
iterationOverride
integer

Override the per-case runs count for this run only.

Required range: 1 <= x <= 10
environmentId
string

Run against one of the suite's attached project environments. Requires suiteId, and must be a member of that suite's environmentIds — otherwise 400 with details.reason: "ENVIRONMENT_NOT_ATTACHED", raised before any case is authored or any server connected.

Omission is meaningful: a suite with no attached environments runs legacy; a suite with exactly ONE attached environment runs against it automatically (the response's environment says which); a suite with several returns 400 with details.reason: "ENVIRONMENT_REQUIRED", naming the candidates.

The environment supplies the closed server set (so serverIds is not required, and is rejected alongside it), and the run is pinned to the revision resolved at launch — if the environment changes in between, the run is rejected with 409 rather than executing against a different configuration.

ephemeralEnvironment
boolean

When true, environmentId may be a project-scoped, non-archived environment that is NOT attached to the suite. The launch never mutates the suite. Absent / false keeps the membership check. Probe GET /environments/capabilities (ephemeralEnvironmentLaunch) before sending — older servers reject the unknown field.

namedHostId
string

Run against ONE host attached to the suite. The platform snapshots that host's current config onto the run and derives the run's server set from it, so a host launch needs no serverIds. Without this, a suite with host attachments runs under the suite's own default host config — the run executes, but the result is attributed to the wrong host.

To run SEVERAL attached hosts, use POST /eval-run-groups rather than N calls here: it is the surface that bounds the fan-out and meters it as one launch.

caseIds
string[]

Narrow the run to these suite cases. The persisted suite is untouched — this filters the run's snapshot only. Every id must belong to the suite; none matching returns 404.

Minimum array length: 1
matchOptionsOverride
object

One-off tool-call match options for this run only, layered over suite defaults and per-case overrides. Does NOT mutate the suite or its cases.

Accepts EITHER the public vocabulary (toolCallOrder: any|in-order|exact, extraToolCalls, arguments) or the internal one (toolCallOrder: ignore|superset|strict, maxExtraToolCalls, argumentMatching). The two are disjoint, so a body can only be one of them; public bodies are normalized server-side.

skillsOverride
enum<string>

The "without skills" arm of an A/B comparison: the run pins NO skills from any channel and is marked skillsExcluded, so the arm is labelled rather than merely empty. Scoped to skill DELIVERY — a pinned plugin's MCP servers stay connected, because which servers an arm connects is the one variable a skills A/B has to hold fixed.

Available options:
exclude
refreshSnapshot
boolean

PERSISTS A SUITE MUTATION. Re-derives and stores the suite's host-config snapshot from this request's server list, so future runs of the suite use it too. Without it a rerun leaves the snapshot frozen, which is what stops newly connected servers from silently contaminating an existing suite. Single-target launches only — it is not accepted on POST /eval-run-groups, where last-writer-wins on a frozen snapshot is never what a fan-out means.

runGroupId
string

A LABEL that groups sibling run rows for display. It has NO quota or launch semantics here: N calls carrying one id are still N independent launches, each metered separately. Grouped-launch behaviour lives on POST /eval-run-groups, which mints the id itself. Echoed back on the 202.

idempotencyKey
string

Write-idempotency key. A repeat call with the same key (same actor and suite) returns the EXISTING run instead of creating and billing a second one. The Idempotency-Key header carries the same value and WINS over this field — it is the transport-level channel unattended clients control, whereas a body key could be shaped by model output.

Maximum string length: 256
sourceHash
string

SHA-256 hex of the suite-file bytes that launched this run. Lowercase, 64 characters. Set by eval run --file; a UI or API launch that did not come from a file omits it.

Pattern: ^[a-f0-9]{64}$

Response

Run created; execution continues in the background.

runId
string
required
suiteId
string
required
status
string
required

The run's status. running on a fresh launch; on a replay (see deduped), the existing run's own status, which may already be terminal.

caseUpsert
object
required

Per-case upsert outcomes for inline tests. Partial failures don't abort the run.

deduped
boolean

Present and true when this request REPLAYED an existing run instead of starting one (an idempotency-key hit, or the short keyless dedupe window). A replayed run is not executed again, so no further credits are spent; read status for what that run actually is. Absent on a fresh launch.

runGroupId
string

Echo of the request's runGroupId, when one was sent. A LABEL only — it groups sibling rows for display and carries no quota or launch semantics.

servers
object[]

The servers the run connects to — explicit or derived from the suite's saved selection. name is present when known (always, on the derived path).

environment
object · null · null

The environment revision this run is pinned to. null on a legacy run that recorded none — always present, so a caller never has to distinguish absent from unpinned.