Open Bias proxy enforces LLM agent rules at runtime

Open Bias, an open source proxy for LLM apps, sits between an application and its model provider to enforce rules from a RULES.md file. The project says it can block, modify, or shadow off-policy behavior in real time without adding default latency.

Open Bias proxy enforces LLM agent rules at runtime

Open Bias is an open source proxy that enforces LLM agent behavior at runtime, according to the project’s Show HN post. The tool sits between an app and an LLM provider and checks requests and responses against rules defined in a repository-hosted RULES.md file.

The project is aimed at developers who want more than prompt-based guardrails. Open Bias ships with a starter RULES.md and, by default, synthesizes an evaluator so users do not need a config file to begin. Developers can point an existing client at the proxy, set the provider API key, and keep using their current app code.

⚡ New to this?

This news is about software that sits between an AI app and the model it calls, then checks whether the AI is following the team’s rules. A proxy is a middle layer, and OpenTelemetry is a tracing system that records what happened so developers can inspect it later. For non-experts, the key point is that this is about enforcing policy during the AI request itself, not just reviewing mistakes afterward.

🦞 OpenClaw angle

If you run self-hosted agents, move your policy checks out of prompts and into a proxy layer like this. Keep rules in a repo file so they can be reviewed, diffed, and deployed with code, then separate fast checks from slower judgment checks. Use blocking for high-risk actions like data deletion or pricing disclosure, and use shadow or next-turn intervention for lower-risk issues where you want logs before hard enforcement.

The project’s example shows a Python OpenAI client pointed at http://localhost:4000/v1 after running pip install openbias and openbias serve. The proxy can sit in front of providers including Anthropic, OpenAI, and Gemini, according to the post.

Open Bias is meant to enforce policy before bad behavior reaches users or production systems. The post says it can intervene on off-policy behavior in real time, rather than only logging or reporting it after the fact.

The authors contrast Open Bias with system prompts and AGENTS.md files, arguing that long instruction lists become less reliable as they grow. They say models treat instructions as context rather than constraints, so more rules in a prompt do not guarantee compliance.

To address that, Open Bias evaluates live traffic against policy and maps the result to enforcement actions. The project says those actions include blocking a request, intervening on the next turn, or shadowing a response for logging while still passing it through.

The proxy uses three hooks on each request: a pre-call hook, a provider call, and a post-call hook. According to the project, pre-call applies pending interventions in microseconds, the provider call is forwarded unmodified, and post-call evaluates the response.

Open Bias also supports multiple policy engines. The project lists a judge engine for compiled rule evaluation, NVIDIA NeMo for content safety and dialog rails, and experimental FSM and LLM engines for state and drift checks.

The authors say different checks do not need the same engine. They describe a setup where a small model can handle fast classification while a judge LLM or NeMo handles more nuanced policy decisions.

The project emphasizes that it is designed not to become a bottleneck. Non-critical violations are evaluated asynchronously and applied on the next turn, while critical violations can be blocked and fixed immediately.

Open Bias also includes OpenTelemetry tracing, a rules compiler, and replay/improve tooling. The team says RULES.md is intended as a shared control surface that lives in the repo, can be reviewed in pull requests, and is versioned alongside code.

The release described in the post is v0.3.0 beta. The project says the proxy layer, judge and NeMo engines, rules compiler, replay tooling, and tracing are working, while FSM and LLM engines remain experimental.

Source: HN Show HN ↗

More from Security News