Skip to main content
The MCPJam API is in preview. It works today and we use it ourselves, but it has not been generally announced and the surface may change — including in breaking ways — while we finish the design. Pin your integration to the behaviors documented on this page, build a tolerant reader, and expect to revisit it. Feedback is very welcome on Discord or GitHub.

Create an API key

Go straight to key management in the hosted app. If you’re signed out, you’ll be asked to sign in and then land right back on the API keys page.
The MCPJam API lets you operate the MCP servers saved in your hosted projects — from CI, scripts, or your own agents — without opening the UI. What you can do with it today:
  • Discover your resources — list your projects, their servers, eval suites, and chat sessions, so every ID the other routes need is self-serve
  • Validate a server — connect, initialize, and capture a capability snapshot
  • Run the doctor — the probe → connect → initialize → capabilities workflow
  • Check OAuth requirements — does this server need an OAuth grant?
  • List tools, prompts, and resources — the server’s MCP primitives
  • Call a tool / render a prompt — execute primitives and get the MCP result back verbatim
  • Read a resource by URI
  • Export a full snapshot — tools, resources, and prompts as one JSON document
  • Manage a project’s hosts — list, read, create (from a built-in template like claude/chatgpt/cursor, or from a full host config), rename, and delete the named model + capability profiles you run chats and eval suites against
  • Manage a project’s environments — list, read, create, edit, archive, and restore the named execution bundles (host + optional server group + optional pinned skills and plugin versions) that eval suites and journeys run against, and preview what one resolves to before launching it
  • List a project’s scenarios — name, access mode, attached servers, and share link — and read one scenario’s settings: model, system prompt, tool-approval policy, and resolved servers
  • Author eval suites — create a runnable suite (name, default model, servers, test cases) synchronously and get a suiteId back; no run is started and no credits are spent
  • Run eval suites asynchronously — create a run from a saved suite, get a 202 + runId immediately, then poll status, per-iteration results (tool calls, token usage, latency), and full traces. Runs appear live in the hosted UI, tagged source: "api"
  • Import OAuth tokens — complete OAuth yourself (e.g. the SDK’s runOAuthLogin) and push the tokens; subsequent calls inject and refresh them server-side
  • Run a headless agent turn — send a message history, get back the assistant reply, the platform operations it invoked, references to any resources it created (eval suites), and token usage; the model is pinned server-side and billed to the project
All endpoints operate on servers you have already added to a project in the hosted inspector. Managing API keys remains UI-only — see Not in the API yet.

Base URL

The API is path-versioned. All routes on this page are relative to the base URL above.

Authentication

Every request needs an MCPJam API key in the Authorization header:

Creating a key

  1. Open Settings → API keys in the hosted app (the card at the top of this page takes you there directly).
  2. Click Create API key, name it, and pick the organization it will act in.
  3. Copy the key (sk_…) immediately — it is shown exactly once and never stored by MCPJam in retrievable form.
Keys can be revoked from the same page at any time. Revocation takes effect immediately.

Scope

A key is bound to one MCPJam organization at creation and acts as you, inside that organization. Per-request authorization still applies: a call only succeeds if the key’s owner can access that project and server. A key can never reach projects outside its organization.
Two deliberate restrictions while in preview:
  • API keys cannot manage API keys. Requests to the key-management surface with an sk_… bearer fail with 403 FORBIDDEN. Create and revoke keys in the UI.
  • Guest sessions cannot use the API. Sign in to create keys.

Keep keys secret

Treat an API key like a password:
  • Load it from an environment variable or secret manager — never commit it to source control or ship it in client-side code.
  • Scope keys narrowly: one key per integration, named so you can tell them apart.
  • Rotate by creating a replacement key, switching traffic, then revoking the old one.
  • If a key leaks, revoke it immediately in Settings → API keys.
If you see 401 UNAUTHORIZED with details.reason: "ORPHANED_KEY", the key is no longer bound to an organization and cannot be used — create a new key from Settings.

Conventions

Requests. Operations on servers are POST with a JSON body (most accept an empty {}). Reads — the catalog listings and eval-run polling — are GET and take their options (cursor, limit, filters) as query parameters. Responses. Three envelope shapes, used consistently: Pagination. Collections are cursor-based. Pass the previous response’s nextCursor as cursor in the next request (body field on POST lists, query parameter on GET lists). Cursors are opaque — don’t parse them. Authoring vs running suites. POST /eval-suites creates a runnable suite (the suite plus its test cases) synchronously and responds 201 with the suiteId, WITHOUT running anything — author once, run later. POST /eval-runs is the async counterpart: it validates and creates a run synchronously, then detaches execution and responds 202 with a runId. Poll GET /eval-runs/{runId} until status is completed, failed, or cancelled. IDs. Every identifier the API takes is discoverable through the API itself: GET /projects lists your projects, GET /projects/{projectId}/servers lists each project’s servers, and GET /projects/{projectId}/eval-suites lists its eval suites. Start from GET /me to confirm which account a key acts as.

Errors

Errors always use the canonical body { code, message, details? }. The code is stable and machine-readable; message is human-readable and may change; details is an optional, unstructured bag. New error codes may be added over time; treat unknown codes as non-retryable failures unless the HTTP status says otherwise.

Rate limits

Each key gets 60 requests per minute sustained, with bursts up to 10. Exceeding it returns 429 RATE_LIMITED with a Retry-After header (in seconds):
Honor Retry-After, add jittered exponential backoff, and expect these limits to be tuned during the preview. Other 429s — the per-minute brakes and daily budgets described with the endpoints below — carry Retry-After whenever the refusal knows when it lifts, which is the ordinary case. Treat it as a hint that is usually there rather than a guarantee: a refusal that reaches you without one still needs your backoff, so never block waiting for a header that may not come.

Endpoints

Each endpoint is fully documented — request and response schemas, examples, and an interactive playground — under Endpoints in the sidebar. Catalog — discover the IDs everything else takes: Project environments — the named execution bundles (one host, optionally a standalone server group, optionally pinned skills and plugin versions) that eval suites and journeys run against. Reads need project membership; every write needs project admin. Writes are revisioned: pass the revision you last read back as expectedRevision, and a stale value returns CONFLICT (409) instead of overwriting a concurrent edit. Archive and restore are explicit sub-actions rather than a DELETE because the row is kept and restore has real semantics: it re-checks the name and drops plugin pins whose version no longer exists. Server diagnostics & primitives (/projects/{projectId}/servers/{serverId}/...): Eval runs (/projects/{projectId}/...): Eval result ingestion (/projects/{projectId}/eval-ingest/...) — how @mcpjam/sdk saves results from eval runs executed outside the platform (local dev, CI). The {projectId} segment accepts the literal default for the key org’s Default project. Most users never call these directly — set MCPJAM_API_KEY and the SDK reporter does: Agent (/projects/{projectId}/agent) — a headless assistant turn for building conversational surfaces (the MCPJam Slack app is the first consumer): Scenarios (/projects/{projectId}/scenarios...) — read-only:
Swarms and User Testing are enabled per organization. Everything in the four sections below is behind that gate, enforced server-side. If it is not on for your organization, the reads generally return empty and every write — creating a persona or journey, launching a run, publishing a scenario, requesting insights — answers 403 with a message naming the feature, not a 404 and not a validation error.Five operations are deliberately outside the gate, because every one of them reduces exposure and spend: cancelling a journey run, unpublishing a scenario, and archiving a persona, journey or swarm (each one’s DELETE). An organization that loses access can still stop what is running and clean up what it authored.Ask us on Discord if you want it turned on.
Swarms — authoring (/projects/{projectId}/personas..., .../journeys..., .../swarms...) — the definitions a swarm run executes. A persona is a reusable synthetic character; a journey is the task you point one at; a swarm is an authoring container holding defaults for the journeys made under it. Nothing here starts anything — launching is POST .../journeys/{journeyId}/runs. Swarms — generation — drafts a model writes for you. Both routes return candidates and persist nothing: feed what you want to keep to the create routes above. That is also why neither accepts an Idempotency-Key — a call with no effect has nothing to de-duplicate, and offering one would imply the drafts are stable across retries when they are not.
The generation routes spend — they run models on your organization’s account. They are metered two ways, and the two mean different waits: a per-minute burst brake (429, retry in seconds) and your plan’s daily budget (429, resets at UTC midnight). Both normally carry Retry-After — honor it rather than guessing, and fall back to your own backoff on the refusal that arrives without one.
Swarms — runs (/projects/{projectId}/journeys/{journeyId}/runs, .../journey-runs/...) — launching a journey and reading what it produced: Three things worth knowing before you write the client:
  • Always send Idempotency-Key on a launch. It spends model credits, so a retry of a dropped response must not run the journey twice. Replaying a key returns the original run with deduped: true and starts no second runner — and consumes no quota. Omitting the header is read as a request that declined to identify itself, and gets a fresh run every time.
  • Check canceled before reporting a failure. A stopped run carries status: "failed", because cancellation is recorded as a marker rather than a status of its own. stale: true is the third case: the runner went silent and the watchdog settled the run.
  • targetId, not hostId, identifies a target. Two environments can resolve to the same host with different servers, so a run can hold two targets that share a hostId.
Launching spends: a run fans out into targets × sessionsPerTarget sessions. Two different 429s can come back — the per-minute burst brake (retry in seconds) and your plan’s daily launch cap (resets at UTC midnight). Both normally carry Retry-After; back off on your own when one does not.
Swarms — insights — what a run revealed. Read these in order of cost: the roll-up and the scorecard are deterministic and free, the findings registry aggregates them, and wave insights run a model. Two arithmetic traps worth naming, because both produce numbers that look plausible and are wrong:
  • Divide by sessionsGraded, never by the session total. Three failures out of four graded sessions, in a run that attempted forty, is 75% — not 7.5%.
  • failedGradingCount is not failCount. A crashed judge is not a regression, and folding the two together makes one look like the other.
And one modelling note: a finding’s status is a lifecycle — new | recurring | regressed | resolved. dismissed is not one of them. Dismissal is the orthogonal dismissedAt, so a finding can be both recurring and dismissed, and hiding one never claims it stopped happening.
Requesting wave insights spends, and draws on the insightsPerDay ledger that is shared with eval-run and user-testing insights — burning it here takes it from there too. Poll the GET rather than re-requesting; force: true deliberately spends a second time.Three refusals, three different meanings: a short 429 is the per-minute brake (wait seconds), a 429 naming insightsPerDay is the daily ledger (its Retry-After, when present, counts to UTC midnight), and a 403 means the feature is not available to your organization — waiting will never help.
User testing (/projects/{projectId}/environments/{environmentId}/scenario, .../user-testing/scenarios/...) — publishing an environment for real visitors, reading what they did, and controlling who can reach it. A scenario is a published environment: one per environment, addressed by its own id once it exists.
Rotating the link does not evict anyone who already used it. It stops the old URL from granting NEW access; the grants already redeemed under it stay valid. That is the single most important thing to know here, because rotation is what you reach for when a link leaks — and on its own it does not undo the leak. Rotate and remove the members you did not intend to have.What DOES take effect at once, ending live sessions rather than waiting for expiry: unpublishing, tightening mode, and removing a member. Those bump the scenario’s accessVersion, and a session minted under an older version stops working. None of them reverse by calling the opposite — a re-invited member gets a new session, not their old one.Read accessVersion off the publish and rebind responses. The update and rotate responses do not carry it.
User testing — what visitors did — the analytics reads, and the model pass over them. Read these in cost order: metrics, usage and signals are deterministic and free; the insights request runs a model. The five analytics shapes above are documented as open objects on purpose. Their upstream projections grow with the product and the SDK types them the same way, so pinning a field list would turn every new metric into a spec violation. Read what you recognize; ignore the rest.
Requesting insights spends, and draws on the same insightsPerDay ledger as eval-run and swarm wave insights — one budget, three producers. Poll the window read rather than re-requesting; force: true deliberately spends a second time, which is why it is never inferred from anything else.
Planning — one read that answers “may I?” before you try:
capabilities is a planning aid, not a gate. Every write still enforces independently, so a true here that races a flag flip costs you a clean 403 rather than an incorrect success — and nothing should consult it instead of trying. It exists because agent surfaces are static: an MCP tool catalog is built with no organization in hand, so the alternative is attempting the write and reading the failure, by which point the agent has usually already told someone what it was about to do.
The same surface is available as an OpenAPI specification for Postman, Swagger UI, or client generation.

Run evals from the API

The shortest useful loop — create a run with one inline test, then poll:
model takes an id from the hosted catalog — provider/name form, e.g. anthropic/claude-haiku-4.5 — and runs on your organization’s credits. Provider-native ids (e.g. claude-sonnet-4-5) are bring-your-own-key: pass the key in modelApiKeys. A model the API can’t execute is rejected at create time with VALIDATION_ERROR; details.hostedModels lists the valid hosted ids for that provider. Reruns are even shorter: { "suiteId": "..." } reruns the suite exactly as configured, connecting the suite’s saved server selection (the 202 response lists the resolved servers). Pass serverIds to override the selection; a suite with no saved selection requires it (VALIDATION_ERROR with details.reason: "NO_SAVED_SERVER_SELECTION" otherwise).

Running against a project environment

A project environment is launchable only through a suite that has it attached. Attach them first with PATCH /projects/{projectId}/eval-suites/{suiteId} and an environmentIds array (send null to detach them all; [] is rejected — use null), then pass environmentId on the run. It must be one of that suite’s attached environments; anything else is a 400 with details.reason: "ENVIRONMENT_NOT_ATTACHED", raised before any case is authored or any server connected. Omitting environmentId is meaningful, and depends on the suite:
  • no attached environments → the run uses the saved server selection, as before;
  • exactly one attached → that environment is used automatically. The 202 response’s environment field says so;
  • several attached → 400 with details.reason: "ENVIRONMENT_REQUIRED", naming the candidates.
The environment supplies the closed server set — including the servers its pinned plugin versions contribute — so serverIds is not required and is rejected alongside it (400), rather than accepted and ignored: honoring both would connect a different set than the run is stamped with. The same goes for a serverIds override on an environment-based suite (details.reason: "ENVIRONMENT_SERVERS_NOT_OVERRIDABLE"). The run is pinned to the environment revision resolved at launch: if the environment changes in between, the run is rejected with CONFLICT (409) rather than executing against a different configuration than the one whose tools were captured. Use GET /projects/{projectId}/environments/{environmentId}/resolve to see what it will connect before launching. Every run records which environment it used. GET .../eval-runs/{runId} (and the run listings) return environment: { id, name, revision } — read from the run’s immutable snapshot, not the suite’s current attachments — or null for a run that used a saved server selection. Per-organization concurrency is capped (default 2 concurrent runs); exceeding it returns 429 with details.reason: "CONCURRENT_RUN_LIMIT" — wait for an active run to finish. If a server answers 401 OAUTH_REQUIRED, complete the OAuth flow yourself (the SDK’s runOAuthLogin handles interactive, headless, and client-credentials flows) and push the result to POST .../oauth/import-tokens once. Subsequent calls inject the stored token and refresh it server-side.

Not in the API yet

To set expectations while in preview, these are not available over the API today (most exist in the hosted inspector UI):
  • Creating or revoking API keys (UI-only by design — see Authentication)
  • Chat and conformance suites
  • Browser-based OAuth flows initiated by the API (use oauth/import-tokens after completing OAuth yourself)
If one of these blocks you, tell us on Discord — it directly shapes what we stabilize first.

Versioning & stability

  • The API is path-versioned (/api/v1). When it reaches general availability, breaking changes will require a new version path.
  • During the preview, breaking changes to v1 may still happen; we’ll note them in the changelog.
  • Additive changes — new endpoints, new optional request fields, new response fields, new error codes — are considered non-breaking and can ship at any time. Write clients that ignore unknown fields.
  • Error codes are stable identifiers; error messages are not.