Create an eval run (async)
Creates a suite run from an existing suiteId (rerun) and/or inline tests, then detaches execution and responds 202 immediately with the runId. Validation and quota errors surface on this request; poll GET /eval-runs/{runId} for progress. The run appears live in the hosted UI Runs tab, tagged source: "api".
A bare suiteId with no inline tests reruns the suite as configured. Per-organization concurrency is capped (default 2 concurrent runs); exceeding it returns 429 with details.reason: "CONCURRENT_RUN_LIMIT".
For a suite with attached project environments, pass environmentId to choose which one the run uses; the 202 echoes the resolved environment triple, and GET /eval-runs/{runId} reports the same triple for the life of the run.
Authorizations
MCPJam API key (sk_…). Create one at Settings → API keys. Guest sessions cannot use the API, and API keys cannot manage other API keys.
Path Parameters
ID of the hosted project that contains the server.
Body
- Option 1
- Option 2
Two valid shapes: suiteId (rerun an existing suite, optionally upserting inline tests into it) or suiteName + tests + serverIds (create a new suite and run it). Inline tests alone — without a suiteId or a suiteName — are rejected with VALIDATION_ERROR.
environmentId requires suiteId: an environment is launchable only through a suite that has it attached (environmentIds, set via PATCH /eval-suites/{suiteId}), so an environment run on a not-yet-created suite could never be satisfied. environmentId and serverIds are mutually exclusive.
Existing suite to rerun. A bare suiteId with no tests reruns the suite exactly as configured.
Name for a new suite. Required (non-empty) when no suiteId is given.
Inline test cases to upsert into the suite before running.
100Servers (by ID) the run connects to. Required when creating a new suite; optional on reruns — when omitted, the run connects the suite's saved server selection (the set its snapshot references). A rerun of a suite with no saved selection is rejected with VALIDATION_ERROR (details.reason: "NO_SAVED_SERVER_SELECTION"). Rejected outright for a suite with attached environments (details.reason: "ENVIRONMENT_SERVERS_NOT_OVERRIDABLE"): the environment supplies a closed set that a server override cannot change, so accepting one would connect a different set than the run is stamped with.
1Optional display names, parallel to serverIds.
When true, skip per-case upsert and rerun the persisted suite. A bare suiteId with no tests is always treated as a rerun.
Optional per-provider model API keys (e.g. { "anthropic": "sk-ant-…" }). Falls back to your organization's configured providers when omitted.
Override the per-case runs count for this run only.
1 <= x <= 10Run against one of the suite's attached project environments. Requires suiteId, and must be a member of that suite's environmentIds — otherwise 400 with details.reason: "ENVIRONMENT_NOT_ATTACHED", raised before any case is authored or any server connected.
Omission is meaningful: a suite with no attached environments runs legacy; a suite with exactly ONE attached environment runs against it automatically (the response's environment says which); a suite with several returns 400 with details.reason: "ENVIRONMENT_REQUIRED", naming the candidates.
The environment supplies the closed server set (so serverIds is not required, and is rejected alongside it), and the run is pinned to the revision resolved at launch — if the environment changes in between, the run is rejected with 409 rather than executing against a different configuration.
When true, environmentId may be a project-scoped, non-archived environment that is NOT attached to the suite. The launch never mutates the suite. Absent / false keeps the membership check. Probe GET /environments/capabilities (ephemeralEnvironmentLaunch) before sending — older servers reject the unknown field.
Run against ONE host attached to the suite. The platform snapshots that host's current config onto the run and derives the run's server set from it, so a host launch needs no serverIds. Without this, a suite with host attachments runs under the suite's own default host config — the run executes, but the result is attributed to the wrong host.
To run SEVERAL attached hosts, use POST /eval-run-groups rather than N calls here: it is the surface that bounds the fan-out and meters it as one launch.
Narrow the run to these suite cases. The persisted suite is untouched — this filters the run's snapshot only. Every id must belong to the suite; none matching returns 404.
1One-off tool-call match options for this run only, layered over suite defaults and per-case overrides. Does NOT mutate the suite or its cases.
Accepts EITHER the public vocabulary (toolCallOrder: any|in-order|exact, extraToolCalls, arguments) or the internal one (toolCallOrder: ignore|superset|strict, maxExtraToolCalls, argumentMatching). The two are disjoint, so a body can only be one of them; public bodies are normalized server-side.
The "without skills" arm of an A/B comparison: the run pins NO skills from any channel and is marked skillsExcluded, so the arm is labelled rather than merely empty. Scoped to skill DELIVERY — a pinned plugin's MCP servers stay connected, because which servers an arm connects is the one variable a skills A/B has to hold fixed.
exclude PERSISTS A SUITE MUTATION. Re-derives and stores the suite's host-config snapshot from this request's server list, so future runs of the suite use it too. Without it a rerun leaves the snapshot frozen, which is what stops newly connected servers from silently contaminating an existing suite. Single-target launches only — it is not accepted on POST /eval-run-groups, where last-writer-wins on a frozen snapshot is never what a fan-out means.
A LABEL that groups sibling run rows for display. It has NO quota or launch semantics here: N calls carrying one id are still N independent launches, each metered separately. Grouped-launch behaviour lives on POST /eval-run-groups, which mints the id itself. Echoed back on the 202.
Write-idempotency key. A repeat call with the same key (same actor and suite) returns the EXISTING run instead of creating and billing a second one. The Idempotency-Key header carries the same value and WINS over this field — it is the transport-level channel unattended clients control, whereas a body key could be shaped by model output.
256SHA-256 hex of the suite-file bytes that launched this run. Lowercase, 64 characters. Set by eval run --file; a UI or API launch that did not come from a file omits it.
^[a-f0-9]{64}$Response
Run created; execution continues in the background.
The run's status. running on a fresh launch; on a replay (see deduped), the existing run's own status, which may already be terminal.
Per-case upsert outcomes for inline tests. Partial failures don't abort the run.
Present and true when this request REPLAYED an existing run instead of starting one (an idempotency-key hit, or the short keyless dedupe window). A replayed run is not executed again, so no further credits are spent; read status for what that run actually is. Absent on a fresh launch.
Echo of the request's runGroupId, when one was sent. A LABEL only — it groups sibling rows for display and carries no quota or launch semantics.
The servers the run connects to — explicit or derived from the suite's saved selection. name is present when known (always, on the derived path).
The environment revision this run is pinned to. null on a legacy run that recorded none — always present, so a caller never has to distinguish absent from unpinned.

