Command Palette

Search for a command to run...

UnylyUnyly
Browse all

Evalbench

FreeNot checked

Offline LLM / agent eval harness with regression gates

GitHubEmbed

About

Offline LLM / agent eval harness with regression gates

README

EVALBENCH

EVALBENCH

Offline LLM / agent eval harness with regression gates

PyPI CI License: COCL 1.0 Suite

AI Agents & LLMOps — build, route, evaluate, and secure agents.

pip install cognis-evalbench
evalbench scan .            # → prioritized findings in seconds

🔎 Example output

Real, reproducible output from the tool — runs offline:

$ evalbench-emit --version
evalbench 2.0.0
$ evalbench-emit --help
usage: evalbench [-h] [--version] {run,demo,gate,types} ...

evalbench — offline eval harness with a regression gate.

positional arguments:
  {run,demo,gate,types}
    run                 evaluate a suite (JSON)
    demo                run the bundled golden-set suite
    gate                compare candidate vs baseline run
    types               list supported assertion types

options:
  -h, --help            show this help message and exit
  --version             show program's version number and exit
$ evalbench-emit demo
evalbench v2.0.0 — eval run: support-bot-golden-set

CASE                      SCORE  RESULT
------------------------------------------------
refund-policy             1.000  PASS
    [ok] icontains     found '30 days'
    [ok] regex         /support@\S+\.\w+/ matched
    [ok] not-contains  "I don't know" absent
    [ok] word-count    words=20 in [8,60]
    [ok] latency       latency=240ms (<= 800.0)
json-extraction           1.000  PASS
    [ok] json-valid    valid JSON
    [ok] json-schema   schema valid
    [ok] json-path     status='shipped'
summary-similarity        0.905  PASS
    [ok] similarity    cosine=0.714 (>= 0.55)
    [ok] contains      found 'Cancel Subscription'
    [ok] length        len=85 in [20,200]
tone-guardrail            1.000  PASS
    [ok] all-of        all-of(regex:P, icontains:P)
    [ok] not-regex     /(?i)stupid|idiot|whatever/ absent
------------------------------------------------
cases: 4/4 passed   pass_rate=1.000   mean_score=0.976

Blocks above are real evalbench output — reproduce them from a clone.

Usage — step by step

  1. Install the harness:

    pip install cognis-evalbench
    
  2. Try the bundled golden set (no files needed), or run your own suite JSON and save the result for later gating:

    evalbench demo
    evalbench run suite.json --save baseline.json
    
  3. Gate a candidate against a baseline. gate accepts saved run results or raw suites and flags score drops beyond --tolerance (add --strict for pass-rate / mean-score regressions):

    evalbench gate baseline.json candidate.json --tolerance 0.02 --format json
    
  4. Read the result. evalbench types lists the assertion types (contains, regex, json-schema, similarity, latency, cost, and more). Exit 0 = success, 1 = case failure / regression, 2 = usage/IO error.

  5. Automate in CI. Evaluate then gate so regressions fail the build:

    evalbench run suite.json --save run.json && evalbench gate baseline.json run.json
    

Contents

Why evalbench?

CI for agents

evalbench is single-purpose, scriptable, and self-hostable: point it at a target, get prioritized results in the format your workflow already speaks (table · JSON · SARIF), gate CI on it, and let agents drive it over MCP.

Features

  • ✅ Load Suite
  • ✅ Run Suite
  • ✅ Compare Baseline
  • ✅ Runs on Linux/macOS/Windows · Docker · devcontainer
  • ✅ Ports in Python, JavaScript, Go, and Rust (ports/)

Quick start

pip install cognis-evalbench
evalbench --version
evalbench scan .                       # scan current project
evalbench scan . --format json         # machine-readable
evalbench scan . --fail-on high        # CI gate (non-zero exit)

Example

$ evalbench scan .
  [HIGH    ] EVA-001  example finding             (./src/app.py)
  [MEDIUM  ] EVA-002  another signal              (./config.yaml)

  2 findings · risk score 5 · 38ms

Architecture

flowchart LR
  IN[agent / A2A traffic] --> P[evalbench<br/>map + analyze]
  P --> OUT[graph + flags]

Use it from any AI stack

evalbench is interoperable with every popular way of using AI:

  • MCP serverevalbench mcp (Claude Desktop, Cursor, Cognis.Studio, uncensored-fleet)
  • OpenAI-compatible / JSON — pipe evalbench scan . --format json into any agent or LLM
  • LangChain · CrewAI · AutoGen · LlamaIndex — wrap the CLI/JSON as a tool in one line
  • CI / scripts — exit codes + SARIF for non-AI pipelines

How it compares

Cognis evalbench promptfoo
Self-hostable, no account varies
Single command, zero config ⚠️
JSON + SARIF for CI varies
MCP-native (AI agents)
Polyglot ports (JS/Go/Rust)
Open license ✅ COCL varies

Built in the spirit of promptfoo / deepeval, re-framed the Cognis way. Missing a credit? Open a PR.

Integrations

Pipes into your stack: SARIF for code-scanning, JSON for anything, an MCP server (evalbench mcp) for AI agents, and a webhook forwarder for SIEM/Slack/Jira. See docs/INTEGRATIONS.md.

Install — every way, every platform

pip install "git+https://github.com/cognis-digital/evalbench.git"    # pip (works today)
pipx install "git+https://github.com/cognis-digital/evalbench.git"   # isolated CLI
uv tool install "git+https://github.com/cognis-digital/evalbench.git" # uv
pip install cognis-evalbench                                          # PyPI (when published)
docker run --rm ghcr.io/cognis-digital/evalbench:latest --help        # Docker
brew install cognis-digital/tap/evalbench                             # Homebrew tap
curl -fsSL https://raw.githubusercontent.com/cognis-digital/evalbench/main/install.sh | sh
Linux macOS Windows Docker Cloud
scripts/setup-linux.sh scripts/setup-macos.sh scripts/setup-windows.ps1 docker run ghcr.io/cognis-digital/evalbench DEPLOY.md (AWS/Azure/GCP/k8s)

Related Cognis tools

  • agentsmith — Config-first scaffolding and orchestration for multi-agent workflows
  • skillhub — Local skill registry and installer for AI agents
  • toolguard — Runtime allowlist and policy for agent tool-calls
  • ragkit — Batteries-included local RAG pipeline — ingest, index, serve
  • memorybank — Portable long-term memory store for agents, exposed over MCP
  • promptpack — Versioned prompt / template registry with A/B and rollbacks

Explore the suite → 🗂️ all 170+ tools · ⭐ awesome-cognis · 🔗 cognis-sources · 🤖 uncensored-fleet · 🧠 engram

Contributing

PRs, new rules, and demo scenarios are welcome under the collaboration-pull model — see CONTRIBUTING.md and SECURITY.md.

⭐ If evalbench saved you time, star it — it genuinely helps others find it.

Interoperability

{} composes with the 300+ tool Cognis suite — JSON in/out and a shared OpenAI-compatible /v1 backbone. See INTEROP.md for the suite map, composition patterns, and reference stacks.

License

Source-available under the Cognis Open Collaboration License (COCL) v1.0 — free for personal, internal-evaluation, research, and educational use; commercial / production use requires a license ([email protected]). See LICENSE.


Cognis Digital · one of 170+ tools in the Cognis Neural Suite · Making Tomorrow Better Today

from github.com/cognis-digital/evalbench

Install Evalbench in Claude Desktop, Claude Code & Cursor

Recommended · one command, every IDE
unyly install evalbench

Installs into Claude Desktop, Claude Code, Cursor & VS Code — handles npx, uvx and build-from-source repos for you.

First time? Get the CLI: curl -fsSL https://unyly.org/install | sh

Or configure manually

Run in your terminal:

claude mcp add evalbench -- uvx --from git+https://github.com/cognis-digital/evalbench cognis-evalbench

Step-by-step: how to install Evalbench

FAQ

Is Evalbench MCP free?

Yes, Evalbench MCP is free — one-click install via Unyly at no cost.

Does Evalbench need an API key?

No, Evalbench runs without API keys or environment variables.

Is Evalbench hosted or self-hosted?

Self-hosted: the server runs locally on your machine via the install command above.

How do I install Evalbench in Claude Desktop, Claude Code or Cursor?

Open Evalbench on unyly.org, pick your client tab (Claude Desktop, Claude Code, Cursor) and press Install — the config is generated automatically, no JSON editing.

Related MCPs

Compare Evalbench with

Not sure what to pick?

Find your stack in 60 seconds

Author?

Embed badge for your README

Browse similar

All ai MCPs