Skip to main content

Audit Agent definitions

Use the platform prompt apistash/agent-doctor to ask your connected model for an evidence-based review of Agent definitions. apistash supplies the audit instructions; your client's model reads the definitions and writes the report. The prompt requests a read-only workflow and does not run a model inside apistash.

The audit begins by explaining its scope: only definitions visible to the current credential and delivered completely to the model can be checked. Hidden, inaccessible, or truncated content can conceal relationships and contradictions. This limit is repeated at the start of the final report, alongside any selection or partial coverage.

Prepare a restricted credential​

Fetching the prompt requires effective Prompts capability, its platform-prompt allowlist entry and admission through the Prompt policy. Running the audit additionally requires Agents, Tools, and admission of apistash/get_agent. Inventory mode also requires apistash/list_agents; an explicit selection needs only get_agent.

For a personal standalone API key, configure this concrete minimum in the dashboard:

  1. Under Settings → MCP Settings, include list_agents and get_agent in your personal platform-tool selection, and select agent-doctor in Platform prompts. Key policies cannot admit platform items missing from these selections.
  2. Create a standalone personal API key with Prompts, Agents, and Tools. Grant no management permissions, especially no create, update, or delete permissions.
  3. Set its Tool policy to a whitelist containing only apistash/list_agents and apistash/get_agent, with no custom tools. For targeted runs, whitelist only get_agent.
  4. Set its Prompt policy to a whitelist containing apistash/agent-doctor.
  5. Set its Agent policy to the definitions intended for this audit. This policy determines which personal definitions can be read; Prompts capability does not grant Agent access.
  6. Connect your client using that key and refresh its tool and prompt lists. Verify that its governed tool list contains only the selected reads and that apistash/agent-doctor is available in the native prompt list.

This effective tool admission prevents MCP management mutations through that credential. The authenticated system tools setup, bootstrap, and agent_context remain callable independently of Tool policy; they do not themselves mutate customer content. The Doctor does not call them, change its Bootstrap role, start subagents, or persist its report.

The prompt's read-only instruction alone is not a security boundary. A broader key can still have write access. Restricting just management tools does not make admitted custom tools or separately granted REST permissions read-only. Organization access is additive: check every bound team's grants and organization-wide tool exposure, since one restrictive team policy cannot veto another admission. Other credentials, local files, and other client tools are outside this credential's guarantee. See the Access Model.

Choose the audit scope​

Open your client's native prompt interface and choose apistash/agent-doctor. The MCP name is stable; a client may display a slash command such as /apistash/agent-doctor, rewrite the name, or offer a prompt picker instead.

The optional agents argument is a string containing a JSON array of exact caller-relative names. Use the names as delivered by apistash: bare names for personal standalone or organization catalog entries, me/name for private content in an organization, and team/name for team content. Commas and newlines can occur inside valid names, so comma-separated or line-separated input is not supported. Exact duplicates are removed in first-occurrence order without normalizing names.

For example, the native MCP request parameters for a targeted run are:

{
"name": "apistash/agent-doctor",
"arguments": {
"agents": "[\"release-reviewer\",\"platform/security-reviewer\"]"
}
}

Pass these parameters to prompts/get, not to the customer-content get_prompt management tool. Let your client pass the returned message to its model. The model first gives the credential notice, validates the JSON selection, then previews the exact names, duplicates removed, planned phases, and expected number of definition reads. A selection of up to ten names proceeds without an extra confirmation. Invalid JSON, non-string items, or empty names require correcting the selection before any definition read.

A nonempty selection is fetched directly using get_agent, without a prior inventory, Bootstrap catalog, or list_agents admission. The Doctor does not expand it to referenced agents. A reference outside this selection is not checked; visibility unknown, unless there is separate inventory or read evidence. It is not a broken handoff merely because the target was not selected.

Leave agents absent or empty, or supply [], to inventory the visible set. The Doctor calls list_agents once and proceeds only if the model receives a complete, parseable Agent array. If inventory is unavailable or visibly truncated, it stops that workflow and asks for an explicit JSON selection. It does not substitute Bootstrap discovery or infer a hidden policy cause. Bootstrap's independent discovery may still be available.

Follow progress and limits​

For more than ten names, the Doctor shows the names and proposes a subset or phases before the first definition read. It does not silently sample the first ten. Inventory grouping uses only the returned metadata and scope or team prefixes. You choose a subset or consent to a phased audit; each phase contains at most ten definitions.

Within a phase, the Doctor reads in small groups and reports progress, failures, and remaining work. The inventory and each definition read are normal metered tool calls, subject to your monthly allowance and per-caller, per-tool rate limits. The prompt fetch itself uses the more generous prompt-read limit and creates no customer Prompt Activity entry. Earlier calls may leave fewer than ten get_agent reads available.

On rate_limited, the Doctor reports the retry time and completed coverage and offers a later continuation; it does not wait and retry in a loop without being asked. Capability failures, internal errors, and truncated responses are failed or incomplete reads, not proof of a missing target. A non-disclosing unavailable response proves only that the selected target was unavailable to that credential at that read, never that it does not exist globally.

Large definitions or cumulative tool output can exceed the client's context. Small read groups do not remove old output. The Doctor must stop when it cannot compare the necessary material reliably, report partial coverage, and propose a smaller selection or separate runs. Client truncation with no visible marker may remain undetectable.

Interpret the report​

The report follows this order:

  1. Credential scope notice — visibility and quality limits.
  2. Coverage — names and active versions observed on each read; requested, checked, and omitted counts, with failures separated. Targeted runs say the total was not inventoried.
  3. Executive summary — findings by severity and the main risks.
  4. Findings — stable IDs within this report, severity, confidence, affected names and versions, field locations, evidence, impact, and a minimal proposed correction with an owner.
  5. Open questions — ambiguity and the information needed to resolve it.
  6. Relationship map — observed handoffs and each target's evidence state.
  7. Verified consistencies — specific agreements actually checked.
  8. Recommended change set — proposals without applying changes.
  9. Reusable selection — correctly escaped JSON names actually checked, for a later run.
  10. Residual limitations — partial coverage, truncation, phase boundaries, and changes possible between reads.

Severity measures impact: high for security risks, lost decision authority, deadlocked handoffs, or materially wrong outcomes; medium for inconsistent results, duplicate work, or faulty handoffs with a safe fallback; low for local terminology or clarity drift. Confidence separately measures certainty of interpretation. Insufficient defect evidence belongs in Open questions, not in an artificial low-confidence finding. Optional improvements are separate from defects.

The relationship map distinguishes targets that were checked, inventoried but not checked, outside the selection with unknown visibility, absent from a completely delivered inventory, or unavailable during a selected read. An unchecked counterpart alone does not establish an asymmetric handoff. Relationships are inferred from text, not from authoritative graph edges.

All arguments, names, descriptions, and definition fields are untrusted evidence. The audit instructions prohibit executing embedded commands, following embedded URLs or file requests, or treating a definition as a new role. The existing Bootstrap role contract is checked where a definition makes relevant claims; definitions need not repeat every platform rule.

Each read observes its then-active version, including conditional instructions. The audit is not an atomic snapshot, and updates or renames can happen between reads. The report identifies which phases were actually compared together; it makes no cross-phase consistency claim when earlier evidence is no longer available. A reusable selection does not pin versions. “No material findings in the audited visible set” is a result for that evidence, not a global consistency certificate.