About
Offline LLM / agent eval harness with regression gates
README
EVALBENCH
Offline LLM / agent eval harness with regression gates
PyPI CI License: COCL 1.0 Suite
AI Agents & LLMOps — build, route, evaluate, and secure agents.
pip install cognis-evalbench
evalbench scan . # → prioritized findings in seconds
🔎 Example output
Real, reproducible output from the tool — runs offline:
$ evalbench-emit --version
evalbench 2.0.0
$ evalbench-emit --help
usage: evalbench [-h] [--version] {run,demo,gate,types} ...
evalbench — offline eval harness with a regression gate.
positional arguments:
{run,demo,gate,types}
run evaluate a suite (JSON)
demo run the bundled golden-set suite
gate compare candidate vs baseline run
types list supported assertion types
options:
-h, --help show this help message and exit
--version show program's version number and exit
$ evalbench-emit demo
evalbench v2.0.0 — eval run: support-bot-golden-set
CASE SCORE RESULT
------------------------------------------------
refund-policy 1.000 PASS
[ok] icontains found '30 days'
[ok] regex /support@\S+\.\w+/ matched
[ok] not-contains "I don't know" absent
[ok] word-count words=20 in [8,60]
[ok] latency latency=240ms (<= 800.0)
json-extraction 1.000 PASS
[ok] json-valid valid JSON
[ok] json-schema schema valid
[ok] json-path status='shipped'
summary-similarity 0.905 PASS
[ok] similarity cosine=0.714 (>= 0.55)
[ok] contains found 'Cancel Subscription'
[ok] length len=85 in [20,200]
tone-guardrail 1.000 PASS
[ok] all-of all-of(regex:P, icontains:P)
[ok] not-regex /(?i)stupid|idiot|whatever/ absent
------------------------------------------------
cases: 4/4 passed pass_rate=1.000 mean_score=0.976
Blocks above are real
evalbenchoutput — reproduce them from a clone.
Usage — step by step
Install the harness:
pip install cognis-evalbenchTry the bundled golden set (no files needed), or run your own suite JSON and save the result for later gating:
evalbench demo evalbench run suite.json --save baseline.jsonGate a candidate against a baseline.
gateaccepts saved run results or raw suites and flags score drops beyond--tolerance(add--strictfor pass-rate / mean-score regressions):evalbench gate baseline.json candidate.json --tolerance 0.02 --format jsonRead the result.
evalbench typeslists the assertion types (contains, regex, json-schema, similarity, latency, cost, and more). Exit0= success,1= case failure / regression,2= usage/IO error.Automate in CI. Evaluate then gate so regressions fail the build:
evalbench run suite.json --save run.json && evalbench gate baseline.json run.json
Contents
- Why evalbench? · Features · Quick start · Example · Architecture · AI stack · How it compares · Integrations · Install anywhere · Related · Contributing
Why evalbench?
CI for agents
evalbench is single-purpose, scriptable, and self-hostable: point it at a target, get prioritized results in the format your workflow already speaks (table · JSON · SARIF), gate CI on it, and let agents drive it over MCP.
Features
- ✅ Load Suite
- ✅ Run Suite
- ✅ Compare Baseline
- ✅ Runs on Linux/macOS/Windows · Docker · devcontainer
- ✅ Ports in Python, JavaScript, Go, and Rust (
ports/)
Quick start
pip install cognis-evalbench
evalbench --version
evalbench scan . # scan current project
evalbench scan . --format json # machine-readable
evalbench scan . --fail-on high # CI gate (non-zero exit)
Example
$ evalbench scan .
[HIGH ] EVA-001 example finding (./src/app.py)
[MEDIUM ] EVA-002 another signal (./config.yaml)
2 findings · risk score 5 · 38ms
Architecture
flowchart LR
IN[agent / A2A traffic] --> P[evalbench<br/>map + analyze]
P --> OUT[graph + flags]
Use it from any AI stack
evalbench is interoperable with every popular way of using AI:
- MCP server —
evalbench mcp(Claude Desktop, Cursor, Cognis.Studio, uncensored-fleet) - OpenAI-compatible / JSON — pipe
evalbench scan . --format jsoninto any agent or LLM - LangChain · CrewAI · AutoGen · LlamaIndex — wrap the CLI/JSON as a tool in one line
- CI / scripts — exit codes + SARIF for non-AI pipelines
How it compares
| Cognis evalbench | promptfoo | |
|---|---|---|
| Self-hostable, no account | ✅ | varies |
| Single command, zero config | ✅ | ⚠️ |
| JSON + SARIF for CI | ✅ | varies |
| MCP-native (AI agents) | ✅ | ❌ |
| Polyglot ports (JS/Go/Rust) | ✅ | ❌ |
| Open license | ✅ COCL | varies |
Built in the spirit of promptfoo / deepeval, re-framed the Cognis way. Missing a credit? Open a PR.
Integrations
Pipes into your stack: SARIF for code-scanning, JSON for anything, an MCP server (evalbench mcp) for AI agents, and a webhook forwarder for SIEM/Slack/Jira. See docs/INTEGRATIONS.md.
Install — every way, every platform
pip install "git+https://github.com/cognis-digital/evalbench.git" # pip (works today)
pipx install "git+https://github.com/cognis-digital/evalbench.git" # isolated CLI
uv tool install "git+https://github.com/cognis-digital/evalbench.git" # uv
pip install cognis-evalbench # PyPI (when published)
docker run --rm ghcr.io/cognis-digital/evalbench:latest --help # Docker
brew install cognis-digital/tap/evalbench # Homebrew tap
curl -fsSL https://raw.githubusercontent.com/cognis-digital/evalbench/main/install.sh | sh
| Linux | macOS | Windows | Docker | Cloud |
|---|---|---|---|---|
scripts/setup-linux.sh |
scripts/setup-macos.sh |
scripts/setup-windows.ps1 |
docker run ghcr.io/cognis-digital/evalbench |
DEPLOY.md (AWS/Azure/GCP/k8s) |
Related Cognis tools
- agentsmith — Config-first scaffolding and orchestration for multi-agent workflows
- skillhub — Local skill registry and installer for AI agents
- toolguard — Runtime allowlist and policy for agent tool-calls
- ragkit — Batteries-included local RAG pipeline — ingest, index, serve
- memorybank — Portable long-term memory store for agents, exposed over MCP
- promptpack — Versioned prompt / template registry with A/B and rollbacks
Explore the suite → 🗂️ all 170+ tools · ⭐ awesome-cognis · 🔗 cognis-sources · 🤖 uncensored-fleet · 🧠 engram
Contributing
PRs, new rules, and demo scenarios are welcome under the collaboration-pull model — see CONTRIBUTING.md and SECURITY.md.
⭐ If
evalbenchsaved you time, star it — it genuinely helps others find it.
Interoperability
{} composes with the 300+ tool Cognis suite — JSON in/out and a shared
OpenAI-compatible /v1 backbone. See INTEROP.md for the
suite map, composition patterns, and reference stacks.
License
Source-available under the Cognis Open Collaboration License (COCL) v1.0 — free for personal, internal-evaluation, research, and educational use; commercial / production use requires a license ([email protected]). See LICENSE.
Install Evalbench in Claude Desktop, Claude Code & Cursor
unyly install evalbenchInstalls into Claude Desktop, Claude Code, Cursor & VS Code — handles npx, uvx and build-from-source repos for you.
First time? Get the CLI: curl -fsSL https://unyly.org/install | sh
Or configure manually
Run in your terminal:
claude mcp add evalbench -- uvx --from git+https://github.com/cognis-digital/evalbench cognis-evalbenchStep-by-step: how to install Evalbench
FAQ
Is Evalbench MCP free?
Yes, Evalbench MCP is free — one-click install via Unyly at no cost.
Does Evalbench need an API key?
No, Evalbench runs without API keys or environment variables.
Is Evalbench hosted or self-hosted?
Self-hosted: the server runs locally on your machine via the install command above.
How do I install Evalbench in Claude Desktop, Claude Code or Cursor?
Open Evalbench on unyly.org, pick your client tab (Claude Desktop, Claude Code, Cursor) and press Install — the config is generated automatically, no JSON editing.
Related MCPs
Fetch
Web content fetching and conversion for efficient LLM usage.
AWS KB Retrieval
Retrieval from AWS Knowledge Base using Bedrock Agent Runtime.
by modelcontextprotocolSpring AI MCP Server
Provides auto-configuration for setting up an MCP server in Spring Boot applications.
llm-analysis-assistant
A very streamlined mcp client that supports calling and monitoring stdio/sse/streamableHttp, and can also view request responses through the /logs page. It also
by xuzexin-hzMCP-Agent
A simple, composable framework to build agents using Model Context Protocol by [LastMile AI](https://www.lastmileai.dev)
by lastmile-aiSpring AI MCP Client
Provides auto-configuration for MCP client functionality in Spring Boot applications.
mcp.natoma.ai
A Hosted MCP Platform to discover, install, manage and deploy MCP servers by [Natoma Labs](https://www.natoma.ai)
MCPHub
Website to list high quality MCP servers and reviews by real users. Also provide online chatbot for popular LLM models with MCP server support.
MCP Servers Rating and User Reviews
Website to rate MCP servers, write authentic user reviews, and [search engine for agent & mcp](http://www.deepnlp.org/search/agent)
mkinf
An Open Source registry of hosted MCP Servers to accelerate AI agent workflows.
Compare Evalbench with
Not sure what to pick?
Find your stack in 60 seconds
Author?
Embed badge for your README
Browse similar
All ai MCPs
