dsh-eval
DeepSeek Harness 的智能体评估平台:基准 YAML、无头 dsh 编排、基于轨迹的指标、LLM 评判、配对 A/B、免密钥重放和跨 harness 导入。
Agent evaluation platform for DeepSeek Harness: benchmark YAML, headless dsh orchestration, trace-based metrics, LLM judge, paired A/B, keyless replay, and cross-harness import.
怎么安装
dsh plugin add github:hccccc01333/dsh-eval 推荐参考
信息
- 协议
- MIT
- 语言
- TypeScript
- GitHub 星标
- 0
- 月下载
- –
- 最近更新
- 2026-08-14
- 创建于
- 2026-08-14
基础安全检查
- 检查结果
- curated 收录但无 npm 包/安装命令
- 收录来源
- curated:0xsline/awesome-deepseek-harness
- 主题标签
- agent-evaluation, benchmark, deepseek-harness, dsh, dsh-plugin, eval
同类推荐
ruflo
ruvnet/claude-flow
🌊 The original agent meta-harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
sandbase-harness
sandbaseai/sandbase-harness
开源、兼容 CMA 的智能体运行时,适用于任何模型,具有 MCP 工具、沙箱会话、审计、重放和本地控制台。包含基于 stdio MCP 的原生 DeepSeek Harness 包。
modlens
liustack/modlens
为纯文本模型架起视觉桥梁:粘贴图片,输出结构化 JSON 证据(OCR、版面、语义)。
pydantic-ai
pydantic/pydantic-ai
How Python does AI: agents, realtime voice, image generation, embeddings. Every model, every interface, typed end to end.