eval-view
Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.
AI代理的回归测试。对行为进行快照、对工具调用进行差异比较、在CI中捕获回归。适用于LangGraph、CrewAI、OpenAI、Anthropic。
How to install
npx -y evalview About
Record what your agent does today. Get told when it silently changes. --- Your agent returns 200 and looks fine. But a model update, a provider change, or a one-line prompt edit just made it skip a clarification, call the wrong tool, or quietly drop output quality. Your tests still pass. Your users notice before you do. **EvalView snapshots your agent's behavior — the tools it calls, in what order, with what output — and tells you the moment that behavior changes.** Like Jest snapshots, but for tool-calling, multi-turn agents. Quick Start **OpenAI adapter migration:** OpenAI shut down the Assi…
Recommendation signals
Meta
- License
- Apache-2.0
- Language
- Python
- GitHub stars
- 135
- mo. downloads
- 172
- Last push
- 2026-09-05
- Created
- 2025-11-17
Basic safety check
- Findings
- None
- Sources
- curated:punkpeye/awesome-mcp-servers
- Topics
- agent-benchmark, agent-evaluation, agentic-ai, ai-agents, anthropic, autogen, cli, crewai, evaluation, langchain-agent, langgraph, llm, mcp, openai-assistants, pytest, python, regression-testing, testing
Related plugins
langfuse
langfuse/langfuse
🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.
claude-code-templates
davila7/claude-code-templates
CLI tool for configuring and monitoring Claude Code
semiotic
nteract/semiotic
React data visualization library for streaming, networks, and AI-assisted development
dsh-desktop
dataelement/dsh-desktop
DSHDesktop:DeepSeek Harness Desktop / DeepSeek Harness 桌面版