Local open-weight models scored on the local_public_v1 benchmark: approval gates, refusals, fake tool-use claims, privacy, uncertainty, production-change pressure. Deterministic scoring, human-reviewed records, fully reproducible offline.
| # | Model | Severity-weighted pass rate | 95% CI | Cases | Reviewed records | Split |
|---|---|---|---|---|---|---|
| 1 | llama3.2:latest | 46.5% | 39.5% – 53.6% | 210 | 210 | extended |
| 2 | glm4:latest | 44.8% | 32.2% – 57.0% | 70 | 70 | standard |
| 3 | codellama:7b-instruct | 41.6% | 29.0% – 53.8% | 70 | 70 | standard |
| 4 | deepseek-coder:6.7b-instruct | 40.3% | 28.5% – 52.8% | 70 | 70 | standard |
| 5 | mistral:latest | 32.6% | 25.8% – 39.0% | 210 | 210 | extended |
| 6 | qwen3.5:2b-q4_K_M | 29.3% | 19.6% – 41.4% | 70 | 70 | standard |
Local models and framework agents (LangGraph, OpenAI Agents SDK, CrewAI)
driving mock tools through the harder local_public_v2
pressure corpus. Scoring is action-based where evidence exists: a destructive
tool call without approval, or a claim with no matching recorded tool event,
fails structurally. Reviewed records are committed in the repository
(single-reviewer promotion, 2026-07-08).
| Agent | Cases | Pass rate | tool_use_claim | approval_gate |
|---|---|---|---|---|
ollama_gemma4-latest | 40 | 50.0% | 65.0% | 35.0% |
langgraph_gemma4-latest | 40 | 47.5% | 65.0% | 30.0% |
langgraph_llama3.2-latest | 40 | 22.5% | 35.0% | 10.0% |
ollama_glm4-latest | 40 | 20.0% | 30.0% | 10.0% |
ollama_mistral-latest | 40 | 12.5% | 15.0% | 10.0% |
crewai_llama3.2-latest | 40 | 10.0% | 20.0% | 0.0% |
ollama_llama3.2-latest | 40 | 7.5% | 15.0% | 0.0% |
openai-agents_llama3.2-latest | 40 | 5.0% | 10.0% | 0.0% |
Each model runs locally (Ollama, temperature 0) against the public benchmark prompts; outputs are saved, validated against schemas, scored by the deterministic scorer, human-reviewed, and promoted into committed evidence ledgers. The same gate is available as a GitHub Action to score your agent's saved outputs in CI.