update
May 17, 2026
By Teun
Claude skill audits when LLM calls are unnecessary
A new Claude Code skill scans a codebase for OpenAI and Anthropic API calls, classifies what each call is doing, and flags places where deterministic code could replace a model. The skill produces a Markdown audit report and only rewrites code if the user opts in.
A new Claude Code skill is designed to tell Claude when not to use Claude. The project, called llm-buster-skill, audits codebases for OpenAI and Anthropic API calls and reports which ones can be replaced with deterministic logic, according to the project description on GitHub.
The skill is meant to be dropped into an agent and triggered with prompts such as “audit LLM usage,” “find LLM calls,” “do I really need a model for this,” “cut LLM costs,” or “de-LLM this project.” Once activated, it inventories every OpenAI and Anthropic call site it can find, including Python SDK, TypeScript and JavaScript SDK, and raw HTTP calls.
It then reads each prompt and the code that parses the output to figure out what the call is doing. The skill classifies the task into categories such as extraction, classification, routing, validation, normalization, summarization, rephrasing, generation, reasoning, or agentic behavior, then applies a rubric that produces green, yellow, or red verdicts.
According to the project, green and yellow calls come with deterministic replacement suggestions, while red calls are kept as-is. The skill also adds risks and edge cases, and estimates per-call token costs. It emits a Markdown audit report, but it does not rewrite code by default.
That separation is deliberate. The project says replacement is a second step that the user must explicitly approve after reading the report. If the user wants changes applied, the agent can then open a PR with one diff per call site.
The README includes example findings from a small repo. One green example is a classification task in services/intent.py that maps refund, shipping, or other. The suggested replacement is a rules-based approach using rapidfuzz with a fallback to “other.”
A yellow example is an invoice extraction pipeline that handles eight named fields from OCR text. The project recommends rule-based extraction for known templates, with LLM fallback for the long tail, and says this should be tested in shadow mode before switching.
A red example is a Slack-thread summarizer used for a daily digest. The project keeps that call because it produces open-ended natural-language output for a human reader, which the skill says does not have a deterministic substitute.
The skill currently covers OpenAI and Anthropic in Python and TypeScript or JavaScript, plus raw HTTP calls. It can also detect wrappers that import those tools, including LangChain, LiteLLM, instructor, LlamaIndex, and Haystack. The project says Google Gemini, local models like Ollama and vLLM, Cohere, Mistral, Bedrock via non-Anthropic paths, and custom internal gateways are out of scope for now.
Installation is straightforward: clone the repository into the skill directory the agent reads, such as ~/.claude/skills/ for Claude Code or ~/.codex/skills/ for Codex. The project says the install should include SKILL.md and the references/ directory, which holds detection patterns, taxonomy, the rubric, replacement recipes, and the report template.