dsh-eval
DeepSeek Harness 的智能体评估平台:基准 YAML、无头 dsh 编排、基于轨迹的指标、LLM 评判、配对 A/B、免密钥重放和跨 harness 导入。
Agent evaluation platform for DeepSeek Harness: benchmark YAML, headless dsh orchestration, trace-based metrics, LLM judge, paired A/B, keyless replay, and cross-harness import.
怎么安装
dsh plugin add github:hccccc01333/dsh-eval 简介
dsh-eval **Agent Evaluation Platform for deepseek-harness.** Run benchmarks against headless dsh profiles, harvest persisted session logs as traces, fold automatic metrics, grade task success and tool selection, and report or compare runs — one benchmark.yaml in, one JSON run + Markdown report out. The dsh ecosystem already has observability and debugging tools (dsh-trace, dsh-tps, dsh-context-doctor). dsh-eval fills the missing slot: **an evaluation platform**. Highlights dsh eval run benchmark.yaml — orchestrate one headless dsh subprocess per case × trial Trace harvesting from persisted ses…
推荐参考
信息
- 协议
- MIT
- 语言
- TypeScript
- GitHub 星标
- 2
- 月下载
- –
- 最近更新
- 2026-08-14
- 创建于
- 2026-08-14
基础安全检查
- 检查结果
- curated 收录但无 npm 包/安装命令
- 收录来源
- curated:0xsline/awesome-deepseek-harness
- 主题标签
- agent-evaluation, benchmark, deepseek-harness, dsh, dsh-plugin, eval
同类推荐
langfuse
langfuse/langfuse
🪢 开源AI工程平台:LLM评估、可观测性、指标、提示管理、游乐场、数据集。与OpenTelemetry、LangChain、OpenAI SDK、LiteLLM等集成。🍊YC W23