← 返回列表

dsh-eval

DeepSeek Harness 开发与运行时 低风险

DeepSeek Harness 的智能体评估平台:基准 YAML、无头 dsh 编排、基于轨迹的指标、LLM 评判、配对 A/B、免密钥重放和跨 harness 导入。

Agent evaluation platform for DeepSeek Harness: benchmark YAML, headless dsh orchestration, trace-based metrics, LLM judge, paired A/B, keyless replay, and cross-harness import.

怎么安装

DeepSeek Harness dsh plugin add github:hccccc01333/dsh-eval

简介

dsh-eval **Agent Evaluation Platform for deepseek-harness.** Run benchmarks against headless dsh profiles, harvest persisted session logs as traces, fold automatic metrics, grade task success and tool selection, and report or compare runs — one benchmark.yaml in, one JSON run + Markdown report out. The dsh ecosystem already has observability and debugging tools (dsh-trace, dsh-tps, dsh-context-doctor). dsh-eval fills the missing slot: **an evaluation platform**. Highlights dsh eval run benchmark.yaml — orchestrate one headless dsh subprocess per case × trial Trace harvesting from persisted ses…

推荐参考

37 工具本身 · 参考星标、下载量、最近更新、安全检查和文档情况
– 用户关注 · 参考最近的查看、安装命令复制和外链访问
37 推荐程度
0访问
0独立访客
0复制安装命令
0下载点击
0外链跳转

信息

协议
MIT
语言
TypeScript
GitHub 星标
2
月下载
–
最近更新
2026-08-14
创建于
2026-08-14

链接

GitHub ↗ 报告问题 ↗

基础安全检查

检查结果
curated 收录但无 npm 包/安装命令
收录来源
curated:0xsline/awesome-deepseek-harness
主题标签
agent-evaluation, benchmark, deepseek-harness, dsh, dsh-plugin, eval

同类推荐

MCP 服务器
推荐100

langfuse

langfuse/langfuse

🪢 开源AI工程平台:LLM评估、可观测性、指标、提示管理、游乐场、数据集。与OpenTelemetry、LangChain、OpenAI SDK、LiteLLM等集成。🍊YC W23

☆ 35.0K ↓ 6.7M 开发与运行时 ↗
Claude Code
推荐92

claude-code-templates

davila7/claude-code-templates

用于配置和监控Claude Code的CLI工具。

☆ 31.7K ↓ 10.7K 开发与运行时 ↗
MCP 服务器
推荐86

semiotic

nteract/semiotic

用于流处理、网络和AI辅助开发的React数据可视化库

☆ 2.7K ↓ 25.0K 开发与运行时 ↗
DeepSeek Harness 精选
推荐84

dsh-desktop

dataelement/dsh-desktop

DeepSeek Harness 桌面版

☆ 9.1K ↓ 930 开发与运行时 ↗