Security for teams of AI agents

Each agent passed review. Together, they leaked the key.

Your engineers run Claude Code, Codex, Cursor and more on one laptop, all sharing the same files. Sherlock finds what those agents can do as a group: the cheapest attack, step by step, or a certificate that none exists.

Runs locallyReads settings, never secretsEvery PASS is certified
attack path · ssh key → internet● live
None of the three agents could do this alone.cost: 0 hacked · 0 approvals

See it work · 36 seconds

Watch Sherlock check a multi-agent app

A support bot written in LangGraph. Sherlock reads the graph, writes the safety question as a formula, and a SAT solver answers it. Toggle the fixes, or let it play.

Live demo · a support bot we built (web search, customer records, email)

Press play for a one-minute walkthrough, or toggle the fixes yourself.

    Computed live in your browser by the same checker as above, from the LangGraph code exported with export(app). Cost = agents an attacker must hack + times you must click Allow.

    Rembrandt's The Night Watch: a militia company with muskets, pikes and a drum, led out by two officers in a pool of light

    On watch

    Every guard checks their own post. Nobody watches the company.

    Each agent is reviewed alone. Sherlock checks how they move together.

    Rembrandt van Rijn, The Night Watch, 1642 · Rijksmuseum, Amsterdam · Public domain

    Scan your Mac

    What can the agents on your laptop do together?

    Three steps, about a minute. The script copies its findings to your clipboard, and this page checks them in your browser. Nothing is uploaded.

    sherlock — scan
    1. Copy the scanner

      One command. It runs scan.py, a 1511-line Python script with no dependencies. It finds every AI agent it can on your Mac and reads the approval settings of 15 of them: Claude Code (per project, with every approval you saved), Codex, GitHub Copilot in VS Code and its CLI, Gemini CLI, Cline, Aider, Continue, OpenCode, Goose, Zed, Kiro, Amazon Q, Devin and their MCP servers. Agents whose settings it cannot read are modeled as unrestricted, and checks whether folders like ~/.ssh exist. It never opens a secret.

      curl -fsSL https://raw.githubusercontent.com/Ning0Luo/sherlock/main/scan.py | python3 -
      Read the script first
    2. Paste it into Terminal

      Open Terminal, paste and press Return. When it prints “Copied”, the results are on your clipboard. It only says that when the scan worked.

      On Linux there is no clipboard tool, so it prints the results instead: add > scan.json to the end and choose the file in step 3.

    3. Paste the results here

    Certified

    When Sherlock says safe, it hands you the certificate

    Scanners and red-team tools can only say they found nothing. Sherlock proves it. Every rule it passes comes with a certificate: evidence that no attack exists within your budget, which you can check yourself without trusting Sherlock.

    • FAILYou get the attack, step by step, with the exact number of hacked agents and approvals it needs.
    • PASSYou get a certificate. For each worst case, it gives a set of facts that holds everything the attacker starts with, is closed under every way influence spreads, and contains no violation. If such a set exists, the attack cannot happen.
    • CHECKverify.py is 152 lines of plain Python. It does no searching; it only checks the evidence. Change one fact and it says INVALID. In our tests it rejected every forged and every tampered certificate.

    A certificate covers the permissions Sherlock read and the attacker it models: someone who writes web pages, hacks up to the agents in your budget and gets up to that many prompts approved. The command-line tool also writes SAT-solver proofs, checked by an independent proof checker.

    Sherlock · sherlock-cert/1Certificate of safety
    Rule
    Web content never rewrites Claude Code's settings
    Holds against
    every choice of 0 hacked agents and 0 approvals
    Setup
    4 agents, 15 places, as scanned
    Evidence
    one closed, violation-free set of facts per worst case
    $ pbpaste | python3 verify.py
    VALID: "Claude Code settings not attacker-controlled" holds for every way to hack 0 agent(s) and approve 0 prompt(s).
    Checked 1 worst case(s) on 4 agents and 15 resources.

    Works with LangGraph

    Check a multi-agent app before you ship it

    Sherlock reads a compiled LangGraph graph without running it: its nodes, the tools each one can call, the state they share, which nodes read which keys, and which tools wait for a person to approve.

    Three lines
    from integrations.langgraph import check
    print(check(app, rules,
      secrets={"data:crm": "customer_pii"}))

    Read automatically, without running the app: nodes, ToolNodes and subgraphs; tools called from plain node code; input_schema; interrupt_before and interrupt() as your approval; the app's input as attacker-written; and the long-term Store. Unknown tools are assumed to do anything.

    Mark what you trust on the node itself: metadata={"sherlock": {"trusted": True, "strips": ["attacker"]}}.

    Scales with SAT: proves a 128-agent app safe in about a second, where trying every combination gives up past 20.

    Live demo

    Watch Sherlock check a LangGraph app, with the formula and the solver's work shown step by step: see the demo at the top of the page ↑

    Why now

    Attackers already use one agent to reach another

    An ordinary program does what its code says. An AI agent may follow instructions hidden in a web page, an issue or a README. Take over one honest agent with text, and the files it shares lead to the rest.

    AUG 2025 · NX “S1NGULARITY”

    Malware drafted the AI CLIs already installed

    A compromised npm package ran the developer's own Claude, Gemini and Q command-line tools and asked them to hunt for secrets and wallets.

    AUG 2025 · CVE-2025-54135

    One write to a config file ran code

    A prompt injection could make Cursor's agent write its own MCP config file, which then ran attacker commands. Control files bridge agents.

    MAY 2025 · GITHUB MCP

    A public issue leaked private repositories

    An agent with access to public and private repos read a malicious issue and copied private data into a public pull request.

    How it works

    Rules in. Attacks or proofs out.

    1. Write the rules you care about

      never ssh -> internet
      tolerate 1 agent, 0 approvals

      Each rule says how many hacked agents and Allow clicks it must survive.

    2. Read every agent's rights

      The scanner reads the permission settings of every agent on the machine. Structure only, never secret values, and nothing leaves the machine.

    3. Solve it as one puzzle

      Rules and rights compile to SAT. A solution is a concrete attack, replayed by a simulator. No solution comes with a certificate that a separate checker verifies.

    4. Fix the cheapest gap

      Each failure names the fewest hacked agents and approvals an attacker needs. Change one setting, re-run in two seconds, and watch the cost rise.

    First results

    Nine rules on a real developer Mac

    0/9rules hold as the agents shipped
    0/9after locking down each agent on its own
    9/9after hardening them together, each with a verified proof
    645test cases where SAT and exhaustive search agree; 647 certificates verified

    Pilot program

    Rolling out coding agents to your engineers?

    Run Sherlock across your developer fleet: one policy file, a scan on every laptop, and a CI gate that blocks a new agent or plugin when it opens a path to your secrets.

    Start with your own Mac →

    Ning Luo · University of Illinois Urbana-Champaign