Claude skill audits when LLM calls are unnecessary

A new Claude Code skill scans a codebase for OpenAI and Anthropic API calls, classifies what each call is doing, and flags places where deterministic code could replace a model. The skill produces a Markdown audit report and only rewrites code if the user opts in.

Claude skill audits when LLM calls are unnecessary

A new Claude Code skill is designed to tell Claude when not to use Claude. The project, called llm-buster-skill, audits codebases for OpenAI and Anthropic API calls and reports which ones can be replaced with deterministic logic, according to the project description on GitHub.

The skill is meant to be dropped into an agent and triggered with prompts such as “audit LLM usage,” “find LLM calls,” “do I really need a model for this,” “cut LLM costs,” or “de-LLM this project.” Once activated, it inventories every OpenAI and Anthropic call site it can find, including Python SDK, TypeScript and JavaScript SDK, and raw HTTP calls.

⚡ New to this?

This is a tool for finding AI calls that do not actually need a model. In plain terms, it checks whether a code path is doing something simple, like classification or extraction, that could be done with regular code instead of paying for an LLM, or large language model. That matters because fewer model calls can mean lower cost, less latency, and simpler systems.

🦞 OpenClaw angle

If you run self-hosted agents, add a similar audit step before you scale traffic or add more prompts. First, inventory every model call in your workflow and tag whether it is extraction, classification, summarization, or open-ended generation. Then separate the safe deterministic cases from the ones that truly need a model, and keep the replacement behind an explicit opt-in step so you can review behavior before changing production code.

It then reads each prompt and the code that parses the output to figure out what the call is doing. The skill classifies the task into categories such as extraction, classification, routing, validation, normalization, summarization, rephrasing, generation, reasoning, or agentic behavior, then applies a rubric that produces green, yellow, or red verdicts.

According to the project, green and yellow calls come with deterministic replacement suggestions, while red calls are kept as-is. The skill also adds risks and edge cases, and estimates per-call token costs. It emits a Markdown audit report, but it does not rewrite code by default.

That separation is deliberate. The project says replacement is a second step that the user must explicitly approve after reading the report. If the user wants changes applied, the agent can then open a PR with one diff per call site.

The README includes example findings from a small repo. One green example is a classification task in services/intent.py that maps refund, shipping, or other. The suggested replacement is a rules-based approach using rapidfuzz with a fallback to “other.”

A yellow example is an invoice extraction pipeline that handles eight named fields from OCR text. The project recommends rule-based extraction for known templates, with LLM fallback for the long tail, and says this should be tested in shadow mode before switching.

A red example is a Slack-thread summarizer used for a daily digest. The project keeps that call because it produces open-ended natural-language output for a human reader, which the skill says does not have a deterministic substitute.

The skill currently covers OpenAI and Anthropic in Python and TypeScript or JavaScript, plus raw HTTP calls. It can also detect wrappers that import those tools, including LangChain, LiteLLM, instructor, LlamaIndex, and Haystack. The project says Google Gemini, local models like Ollama and vLLM, Cohere, Mistral, Bedrock via non-Anthropic paths, and custom internal gateways are out of scope for now.

Installation is straightforward: clone the repository into the skill directory the agent reads, such as ~/.claude/skills/ for Claude Code or ~/.codex/skills/ for Codex. The project says the install should include SKILL.md and the references/ directory, which holds detection patterns, taxonomy, the rubric, replacement recipes, and the report template.

Source: HN Show HN ↗

More from OpenClaw News