eval-view
AI代理的回归测试。对行为进行快照、对工具调用进行差异比较、在CI中捕获回归。适用于LangGraph、CrewAI、OpenAI、Anthropic。
Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.
怎么安装
npx -y evalview 简介
Record what your agent does today. Get told when it silently changes. --- Your agent returns 200 and looks fine. But a model update, a provider change, or a one-line prompt edit just made it skip a clarification, call the wrong tool, or quietly drop output quality. Your tests still pass. Your users notice before you do. **EvalView snapshots your agent's behavior — the tools it calls, in what order, with what output — and tells you the moment that behavior changes.** Like Jest snapshots, but for tool-calling, multi-turn agents. Quick Start **OpenAI adapter migration:** OpenAI shut down the Assi…
推荐参考
信息
- 协议
- Apache-2.0
- 语言
- Python
- GitHub 星标
- 135
- 月下载
- 172
- 最近更新
- 2026-09-05
- 创建于
- 2025-11-17
基础安全检查
- 检查结果
- 无
- 收录来源
- curated:punkpeye/awesome-mcp-servers
- 主题标签
- agent-benchmark, agent-evaluation, agentic-ai, ai-agents, anthropic, autogen, cli, crewai, evaluation, langchain-agent, langgraph, llm, mcp, openai-assistants, pytest, python, regression-testing, testing
同类推荐
langfuse
langfuse/langfuse
🪢 开源AI工程平台:LLM评估、可观测性、指标、提示管理、游乐场、数据集。与OpenTelemetry、LangChain、OpenAI SDK、LiteLLM等集成。🍊YC W23