MCP Server Guide¶
Use LLM Council as a Model Context Protocol (MCP) server with Claude Code or Claude Desktop.
Installation¶
Claude Code Setup¶
# Store API key securely
llm-council setup-key
# Add MCP server
claude mcp add llm-council --scope user -- llm-council
Claude Desktop Setup¶
Add to ~/Library/Application Support/Claude/claude_desktop_config.json:
Available Tools¶
consult_council¶
Ask the LLM council a question.
Arguments:
| Argument | Type | Default | Description |
|---|---|---|---|
query |
string | required | Question to ask |
confidence |
string | "high" |
quick, balanced, high, reasoning |
verdict_type |
string | "synthesis" |
synthesis, binary, tie_breaker |
include_details |
boolean | false |
Individual responses + full cost breakdown |
include_dissent |
boolean | false |
Include minority opinions |
evidence |
list | none | Caller-supplied grounding context (#619, ADR-042) |
on_partial |
string | "synthesise" |
What to do when the council doesn't complete: synthesise, error, return_raw |
Every response ends with a one-line Cost & Tokens summary (ADR-011);
include_details=true adds the per-model/per-stage breakdown.
Reading a degraded answer¶
A council run can complete with fewer models than it asked for, and the heading tells you which kind of answer you got. This matters more than it looks: the models that time out are the slowest, which are generally the strongest reasoners, so a short council is not a random sample of the full one — it is skewed toward the faster and weaker members.
| Heading | What it is |
|---|---|
### Chairman's Synthesis |
Full deliberation: every member responded, peer review ran, the chairman synthesised |
### Chairman's Synthesis — N of M models |
Real synthesis including peer review; only the membership was short |
### Partial synthesis — N of M models, no peer review |
The run hit its global deadline; the chairman synthesised stage-1 drafts directly, stage 2 never ran |
### Single-model response from <model> — council incomplete (N/M), chairman unavailable |
Not a synthesis: the chairman failed too, so this is one surviving member's raw text |
### Council Failed |
No usable responses |
Any shortfall is disclosed in a > **Note** above the answer, never below
it. Every response also ends with a machine-readable block so automation can
branch without parsing prose:
{"status": "partial", "synthesis_type": "single_model_raw", "models_responded": 2,
"models_requested": 4, "peer_review": false, "tier": "high",
"failed_models": [{"model": "anthropic/claude-opus-5", "status": "timeout"}]}
Treat peer_review: false as the important flag: anonymised peer review is
the mechanism that filters confident-but-wrong answers, and without it you
are reading opinions rather than a deliberated verdict.
include_dissent=true now works in the default verdict_type="synthesis"
mode (previously it was only rendered for binary/tie_breaker, so the
extracted dissent was silently discarded). When nothing is surfaced you get a
Dissent section saying why — an empty section and no section at all mean
different things.
Choosing what a partial council does (on_partial)¶
Labelling helps a human reading the output. It does nothing for an automated
caller that will act on the text either way, so on_partial lets you decide
up front:
| Value | Behaviour |
|---|---|
synthesise (default) |
Answer anyway, with the shortfall in the heading and above the content |
error |
Return a council_incomplete JSON blob instead of an answer — the synthesis text is not included, so there is nothing to accidentally act on |
return_raw |
Return the surviving members' responses attributed individually, with no chairman synthesis over the top |
Use error when you asked for a full council and would rather retry than act
on a short one — for a gate, or any automated decision. Use return_raw when
you want to judge the disagreement yourself: an unsynthesised set of attributed
answers shows the divergence a synthesis would have ironed out.
An unrecognised value is rejected with invalid_on_partial rather than
silently defaulting. (confidence does silently fall back to high; that
would be the wrong choice here, since a typo would turn a strictness request
into permissiveness.)
Grounding the council with your own context (evidence). If your client
already has retrieval — web search, a RAG index, repo files — you can hand the
retrieved snippets to the council instead of hoping the models know them.
Each item is a dict: source (required, tool@version-style name), content
(required), optional format (markdown/json/text), evidence_id, and
strength. The items are rendered into the question for every council
member, clearly fenced as data (models are instructed not to follow
instructions inside evidence bodies), under the same per-tier budget as
verify's evidence (quick 1.5K / balanced 6K / high & reasoning 10K chars —
whole items are dropped when over budget, never truncated mid-string, and the
response tells you which). Two things to know:
consult_councilhas no pass/fail gate, sostrength="blocking"has no meaning here — such items are downgraded to informational with an explicit note in the response. Useverify()when you want gate semantics.- No
evidence⇒ the query is sent byte-identically as before.
Set MCP_TIMEOUT for high/reasoning — these numbers are for consult_council only
These tiers exceed many clients' default transport timeout (~60s). Set
MCP_TIMEOUT (milliseconds) in your client config — e.g. 180000 for
high, 600000 for reasoning — or the client will drop the connection
while the council deliberates.
verify needs roughly double these values. Its global deadline is
tier_deadline × 2.0, so a reasoning verify can run to 1200s and these
numbers would cut it off at the halfway mark. See
the verify guide's tier table.
MCP_TIMEOUT is the client transport budget. Server-side, each tier also
budgets its own stages: Stage 1 gets the tier's per-model timeout, and
Stage 2 (peer review) and Stage 3 (chairman synthesis) get
max(per-model, 120s) — 120s at quick/balanced/high, 300s at reasoning.
Set LLM_COUNCIL_TIMEOUT_MULTIPLIER to scale all of them together if your
chairman model is slow (e.g. 3 triples every tier budget). Before
#648 both of those
stages were pinned at 120s regardless of tier, which cost reasoning-tier
runs their synthesis.
Example:
Use consult_council with confidence="balanced" to ask:
"What are the trade-offs between REST and GraphQL?"
verify¶
Multi-model verification of code, documents, or any work product with a
machine-actionable verdict — the CI-gate surface. See the
Verification & CI Gating guide for tiers, unclear_reason
routing, calibrated confidence, screening, and evidence injection.
Arguments:
| Argument | Type | Default | Description |
|---|---|---|---|
snapshot_id |
string | required | Git commit SHA (≥7 hex chars) |
target_paths |
list | none | Files/dirs to verify (scope to the change) |
tier |
string | "balanced" |
quick, balanced, high, reasoning |
rubric_focus |
string | none | e.g. Security, Performance |
confidence_threshold |
float | 0.7 |
Minimum confidence for PASS |
evidence |
list | none | Upstream tool findings (ADR-042) |
Returns verdict/confidence (raw + calibrated), rubric scores, blocking
issues, unclear_reason on UNCLEAR, and the transcript location. Under
LLM_COUNCIL_STRUCTURED_FINDINGS (ADR-051) it also returns a typed findings
array and computes the verdict from it — see the verify guide's
Response fields and structured-findings sections.
audit¶
Retrieve the persisted transcript for a past verification (by
verification_id) — the audit trail behind every verdict.
council_health_check¶
Verify the council is ready.
Parameters:
tier(default"high"): report readiness for the tier a real run would use. Mirrorsconsult_council's resolution, including its fallback tohighfor an unrecognised value.deep(defaulttruesince #660): probe the configured chairman model as well as general API reachability. Costs one small chairman call (~2-3s) on top of the lite ping. Passdeep=falsefor the old cheap-ping-only behaviour. Skipped automatically when general connectivity has already failed — a chairman probe adds nothing then, and billing for one during an outage is the wrong move.
Returns:
api_key_configured: Whether key is setkey_source: Where key came fromdefault_tier: The tier whose models are reported belowcouncil_size/models: The models a realconsult_councilrun would use — resolved from the tier pool, which is what the council actually runsconfigured_council_models/config_warnings: Present only when the flatcouncil.modelslist disagrees with the resolved tier pool, so the two cannot diverge silentlyapi_connectivity.probe_scope:connectivity_onlyfor the default probe, with acaveatnaming what it does not coverchairman_connectivity: Present only withdeep=trueestimated_duration: Per-tier server budget, derived from each tier's configuredtimeout_seconds(and anyLLM_COUNCIL_TIMEOUT_MULTIPLIER) rather than hand-written prose — an upper bound, not a typical latencyready: Whether council is operational — seeready_scopefor what that claim coversready_scope:connectivity_only(default) orchairman_probed(deep=true). Sits besidereadydeliberately: a caveat nested insideapi_connectivityis one a caller has to go looking for
Read ready_scope before trusting ready
ready: true under ready_scope: "connectivity_only" means "the API is reachable", not "the council will complete". Those diverge during a chairman outage — stage-3 synthesis is a single point of failure, so a healthy API can still yield runs with no verdict, which is what happened for ~an hour on 2026-07-16 (#596).
Since #660 the default is ready_scope: "chairman_probed", where ready does account for the chairman. You only get the narrower claim if you asked for it with deep=false, or if connectivity failed before the chairman could be probed. Either way the field tells you which claim you're holding — never infer it from the deep argument you passed, since a skipped probe reports the narrower scope.
Jury Mode¶
For binary decisions:
Use consult_council with verdict_type="binary" to ask:
"Should we approve this architectural change?"
Returns: