Claude Opus 4.7 outscored Kimi K2.6 in a head-to-head build test for a workflow orchestration API, according to the review published by OpenClawsome.com. The two models were given the same 1,042-line FlowGraph specification and asked to build the project from scratch in empty directories.
The result was 91/100 for Claude Opus 4.7 and 68/100 for Kimi K2.6. Kimi reached about 75% of Claude’s score while costing roughly 19% as much, but the gap showed up in the parts of the system that are hardest to get right: lease handling, scheduling, and live event streaming.
FlowGraph is described as a persistent workflow orchestration API. In practical terms, that means it manages multi-step jobs with dependencies, retries, worker claims, lease expiry recovery, pause/resume/cancel actions, and server-sent events, or SSE, for event streaming.
The test used the same prompt for both models: read the spec, build the project in the current directory, create all code and tests, install dependencies, run the suite, fix failures, and leave a runnable project behind. Claude ran in high thinking mode. Kimi ran in thinking mode.
Both models produced the broad project shape the spec called for. The builds included Prisma with SQLite as the source of truth, Hono routes for workflow definitions, runs, worker actions, events, health, and metrics, conditional updateMany logic for step claiming, retry and lease-expiry scheduling, a RunEvent audit table, and README files with setup notes.
The first pass looked good because both test suites passed. Claude ran 31 tests across six files and Kimi ran 20 tests in one file, with no failures reported. But the reviewers did not stop there. They reviewed the code directly and reproduced edge cases against isolated SQLite databases that the model-written tests had missed.
Claude Opus 4.7 had one confirmed bug. In its recovery path, if two expired leases were processed in the same pass, a step that had just been blocked by a failed run could be moved back to waiting_retry because the update in handleLeaseExpiry() only checked the step id. The review team reproduced this with two expired running steps in one run: one with maxAttempts = 1 and another with maxAttempts = 2. After recovery, the second step should have stayed blocked, but instead it could be claimed again.
The review also flagged two smaller risks in Claude’s implementation. Its claim path only scanned maxClaims * 10 candidates, which could skip valid work if many early candidates were ineligible. Its SSE stream also treated an unknown cursor as “replay everything,” which the spec did not clearly define.
Kimi K2.6 had six confirmed issues. The biggest was scheduling: the spec says claim order must be global across all eligible steps, sorted by priority descending, then availableAt ascending, then createdAt ascending. Kimi ordered steps inside each run, then iterated runs in database order, which let a lower-priority step win before a higher-priority one in another run.
Its event stream also did not become live after replay. The route replayed stored events and then only kept a timer alive. Although src/lib/events.ts defined an emitAndBroadcast function and subscriber map, the stream route never subscribed to new events. The README still claimed live streaming.
The review found another correctness issue in Kimi’s worker actions. The heartbeat endpoint rejected expired leases, but the complete and fail endpoints did not. That meant a worker could let a lease expire, have the system schedule a retry, and still later report success on the expired lease.
There were also API mismatches. If no active workflow version existed and no explicit version was supplied, the spec required a 409 response, but Kimi returned 404. Its validation was narrower than the spec too: it used z.record(z.any()) for input, metadata, and output, which rejects JSON values like arrays, strings, or numbers even though the spec allows arbitrary JSON payloads.
The final issue was in the documented build flow. npm test passed, but npm run build did not produce a working start path because package.json expected npm start to run node dist/index.js. On a clean checkout, that flow was broken.
The article’s author argues the test shows a clear split between surface correctness and deep correctness. Both models could generate the right endpoint set and pass their own tests. Only targeted reproduction exposed the differences in lease recovery, cross-run scheduling, and live streaming behavior.