dsh-eval
Agent evaluation platform for DeepSeek Harness: benchmark YAML, headless dsh orchestration, trace-based metrics, LLM judge, paired A/B, keyless replay, and cross-harness import.
DeepSeek Harness 的智能体评估平台:基准 YAML、无头 dsh 编排、基于轨迹的指标、LLM 评判、配对 A/B、免密钥重放和跨 harness 导入。
How to install
dsh plugin add github:hccccc01333/dsh-eval About
dsh-eval **Agent Evaluation Platform for deepseek-harness.** Run benchmarks against headless dsh profiles, harvest persisted session logs as traces, fold automatic metrics, grade task success and tool selection, and report or compare runs — one benchmark.yaml in, one JSON run + Markdown report out. The dsh ecosystem already has observability and debugging tools (dsh-trace, dsh-tps, dsh-context-doctor). dsh-eval fills the missing slot: **an evaluation platform**. Highlights dsh eval run benchmark.yaml — orchestrate one headless dsh subprocess per case × trial Trace harvesting from persisted ses…
Recommendation signals
Meta
- License
- MIT
- Language
- TypeScript
- GitHub stars
- 2
- mo. downloads
- –
- Last push
- 2026-08-14
- Created
- 2026-08-14
Links
Basic safety check
- Findings
- curated 收录但无 npm 包/安装命令
- Sources
- curated:0xsline/awesome-deepseek-harness
- Topics
- agent-evaluation, benchmark, deepseek-harness, dsh, dsh-plugin, eval
Related plugins
langfuse
langfuse/langfuse
🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.
claude-code-templates
davila7/claude-code-templates
CLI tool for configuring and monitoring Claude Code
semiotic
nteract/semiotic
React data visualization library for streaming, networks, and AI-assisted development
dsh-desktop
dataelement/dsh-desktop
DSHDesktop:DeepSeek Harness Desktop / DeepSeek Harness 桌面版