Skip to content

MCP Server Guide

Use LLM Council as a Model Context Protocol (MCP) server with Claude Code or Claude Desktop.

Installation

pip install "llm-council-core[mcp]"

Claude Code Setup

# Store API key securely
llm-council setup-key

# Add MCP server
claude mcp add llm-council --scope user -- llm-council

Claude Desktop Setup

Add to ~/Library/Application Support/Claude/claude_desktop_config.json:

{
  "mcpServers": {
    "llm-council": {
      "command": "llm-council"
    }
  }
}

Available Tools

consult_council

Ask the LLM council a question.

Arguments:

Argument Type Default Description
query string required Question to ask
confidence string "high" quick, balanced, high, reasoning
verdict_type string "synthesis" synthesis, binary, tie_breaker
include_details boolean false Individual responses + full cost breakdown
include_dissent boolean false Include minority opinions
evidence list none Caller-supplied grounding context (#619, ADR-042)
on_partial string "synthesise" What to do when the council doesn't complete: synthesise, error, return_raw

Every response ends with a one-line Cost & Tokens summary (ADR-011); include_details=true adds the per-model/per-stage breakdown.

Reading a degraded answer

A council run can complete with fewer models than it asked for, and the heading tells you which kind of answer you got. This matters more than it looks: the models that time out are the slowest, which are generally the strongest reasoners, so a short council is not a random sample of the full one — it is skewed toward the faster and weaker members.

Heading What it is
### Chairman's Synthesis Full deliberation: every member responded, peer review ran, the chairman synthesised
### Chairman's Synthesis — N of M models Real synthesis including peer review; only the membership was short
### Partial synthesis — N of M models, no peer review The run hit its global deadline; the chairman synthesised stage-1 drafts directly, stage 2 never ran
### Single-model response from <model> — council incomplete (N/M), chairman unavailable Not a synthesis: the chairman failed too, so this is one surviving member's raw text
### Council Failed No usable responses

Any shortfall is disclosed in a > **Note** above the answer, never below it. Every response also ends with a machine-readable block so automation can branch without parsing prose:

{"status": "partial", "synthesis_type": "single_model_raw", "models_responded": 2,
 "models_requested": 4, "peer_review": false, "tier": "high",
 "failed_models": [{"model": "anthropic/claude-opus-5", "status": "timeout"}]}

Treat peer_review: false as the important flag: anonymised peer review is the mechanism that filters confident-but-wrong answers, and without it you are reading opinions rather than a deliberated verdict.

include_dissent=true now works in the default verdict_type="synthesis" mode (previously it was only rendered for binary/tie_breaker, so the extracted dissent was silently discarded). When nothing is surfaced you get a Dissent section saying why — an empty section and no section at all mean different things.

Choosing what a partial council does (on_partial)

Labelling helps a human reading the output. It does nothing for an automated caller that will act on the text either way, so on_partial lets you decide up front:

Value Behaviour
synthesise (default) Answer anyway, with the shortfall in the heading and above the content
error Return a council_incomplete JSON blob instead of an answer — the synthesis text is not included, so there is nothing to accidentally act on
return_raw Return the surviving members' responses attributed individually, with no chairman synthesis over the top

Use error when you asked for a full council and would rather retry than act on a short one — for a gate, or any automated decision. Use return_raw when you want to judge the disagreement yourself: an unsynthesised set of attributed answers shows the divergence a synthesis would have ironed out.

An unrecognised value is rejected with invalid_on_partial rather than silently defaulting. (confidence does silently fall back to high; that would be the wrong choice here, since a typo would turn a strictness request into permissiveness.)

Grounding the council with your own context (evidence). If your client already has retrieval — web search, a RAG index, repo files — you can hand the retrieved snippets to the council instead of hoping the models know them. Each item is a dict: source (required, tool@version-style name), content (required), optional format (markdown/json/text), evidence_id, and strength. The items are rendered into the question for every council member, clearly fenced as data (models are instructed not to follow instructions inside evidence bodies), under the same per-tier budget as verify's evidence (quick 1.5K / balanced 6K / high & reasoning 10K chars — whole items are dropped when over budget, never truncated mid-string, and the response tells you which). Two things to know:

  • consult_council has no pass/fail gate, so strength="blocking" has no meaning here — such items are downgraded to informational with an explicit note in the response. Use verify() when you want gate semantics.
  • No evidence ⇒ the query is sent byte-identically as before.

Set MCP_TIMEOUT for high/reasoning — these numbers are for consult_council only

These tiers exceed many clients' default transport timeout (~60s). Set MCP_TIMEOUT (milliseconds) in your client config — e.g. 180000 for high, 600000 for reasoning — or the client will drop the connection while the council deliberates.

verify needs roughly double these values. Its global deadline is tier_deadline × 2.0, so a reasoning verify can run to 1200s and these numbers would cut it off at the halfway mark. See the verify guide's tier table.

MCP_TIMEOUT is the client transport budget. Server-side, each tier also budgets its own stages: Stage 1 gets the tier's per-model timeout, and Stage 2 (peer review) and Stage 3 (chairman synthesis) get max(per-model, 120s) — 120s at quick/balanced/high, 300s at reasoning. Set LLM_COUNCIL_TIMEOUT_MULTIPLIER to scale all of them together if your chairman model is slow (e.g. 3 triples every tier budget). Before #648 both of those stages were pinned at 120s regardless of tier, which cost reasoning-tier runs their synthesis.

Example:

Use consult_council with confidence="balanced" to ask:
"What are the trade-offs between REST and GraphQL?"

verify

Multi-model verification of code, documents, or any work product with a machine-actionable verdict — the CI-gate surface. See the Verification & CI Gating guide for tiers, unclear_reason routing, calibrated confidence, screening, and evidence injection.

Arguments:

Argument Type Default Description
snapshot_id string required Git commit SHA (≥7 hex chars)
target_paths list none Files/dirs to verify (scope to the change)
tier string "balanced" quick, balanced, high, reasoning
rubric_focus string none e.g. Security, Performance
confidence_threshold float 0.7 Minimum confidence for PASS
evidence list none Upstream tool findings (ADR-042)

Returns verdict/confidence (raw + calibrated), rubric scores, blocking issues, unclear_reason on UNCLEAR, and the transcript location. Under LLM_COUNCIL_STRUCTURED_FINDINGS (ADR-051) it also returns a typed findings array and computes the verdict from it — see the verify guide's Response fields and structured-findings sections.

audit

Retrieve the persisted transcript for a past verification (by verification_id) — the audit trail behind every verdict.

council_health_check

Verify the council is ready.

Parameters:

  • tier (default "high"): report readiness for the tier a real run would use. Mirrors consult_council's resolution, including its fallback to high for an unrecognised value.
  • deep (default true since #660): probe the configured chairman model as well as general API reachability. Costs one small chairman call (~2-3s) on top of the lite ping. Pass deep=false for the old cheap-ping-only behaviour. Skipped automatically when general connectivity has already failed — a chairman probe adds nothing then, and billing for one during an outage is the wrong move.

Returns:

  • api_key_configured: Whether key is set
  • key_source: Where key came from
  • default_tier: The tier whose models are reported below
  • council_size / models: The models a real consult_council run would use — resolved from the tier pool, which is what the council actually runs
  • configured_council_models / config_warnings: Present only when the flat council.models list disagrees with the resolved tier pool, so the two cannot diverge silently
  • api_connectivity.probe_scope: connectivity_only for the default probe, with a caveat naming what it does not cover
  • chairman_connectivity: Present only with deep=true
  • estimated_duration: Per-tier server budget, derived from each tier's configured timeout_seconds (and any LLM_COUNCIL_TIMEOUT_MULTIPLIER) rather than hand-written prose — an upper bound, not a typical latency
  • ready: Whether council is operational — see ready_scope for what that claim covers
  • ready_scope: connectivity_only (default) or chairman_probed (deep=true). Sits beside ready deliberately: a caveat nested inside api_connectivity is one a caller has to go looking for

Read ready_scope before trusting ready

ready: true under ready_scope: "connectivity_only" means "the API is reachable", not "the council will complete". Those diverge during a chairman outage — stage-3 synthesis is a single point of failure, so a healthy API can still yield runs with no verdict, which is what happened for ~an hour on 2026-07-16 (#596).

Since #660 the default is ready_scope: "chairman_probed", where ready does account for the chairman. You only get the narrower claim if you asked for it with deep=false, or if connectivity failed before the chairman could be probed. Either way the field tells you which claim you're holding — never infer it from the deep argument you passed, since a skipped probe reports the narrower scope.

Jury Mode

For binary decisions:

Use consult_council with verdict_type="binary" to ask:
"Should we approve this architectural change?"

Returns:

{
  "verdict": "approved",
  "confidence": 0.75,
  "rationale": "Council agreed..."
}