Token导航 LogoToken导航TokenDH.com
研究检索只读clawhub未标认证来源可访问clear审计提醒

model-tester模型测试员

Agent Skill

用于辅助测试设计、自动化测试、用例整理和回归验证。它适合让 Agent 编写单元测试、端到端测试、测试计划或根据失败日志定位问题。使用时需要确认项目测试框架、运行命令和夹具数据,避免为了通过测试而改坏真实逻辑;涉及浏览器或外部服务时,应区分本地模拟、测试环境和生产环境。

总安装

11,078

周安装

471

GitHub Stars

公开资料未说明

下载量

3,881
OpenClaw

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

MIT-0

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:model-tester(模型测试员)
来源仓库:https://github.com/nandorocker/model-tester
安装命令:
openclaw skills install model-tester
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 OpenClaw 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

ClawHubOpenClaw
openclaw skills install model-tester

简介

用于根据预定义测试用例验证模型路由、性能和输出质量。model-tester 属于研究检索类 Skill,可作为该场景下的辅助能力补充。

  • 适合在开发阶段验证代理行为、定位问题或回归测试时使用。
  • 通过 clawhub 安装,需确认项目测试框架和运行环境。
  • 建议结合原始 README 核验测试用例格式和执行流程。
  • 使用前请评估是否会触发命令执行或外部服务调用,避免误改生产逻辑。

SKILL.md

name
model-tester
version
1.0.0
description
Test agents or models against predefined test cases to validate model routing, performance, and output quality. Use when: (1) verifying a specific agent or model works correctly, (2) debugging model fallback chains, (3) testing model selection behavior, (4) validating extraction/reasoning/classification across different models, or (5) verifying a model actually got used after routing. Supports --agent, --model, --case parameters with structured JSON output.

Use scripts/model_tester.py to run repeatable test prompts and compare requested vs actual model usage from OpenClaw logs.

Run

From the skill directory (or pass absolute paths):

python3 scripts/model_tester.py --agent menial --case extract-emails
python3 scripts/model_tester.py --model openai/gpt-4.1 --case math-reasoning
python3 scripts/model_tester.py --agent chat --model openai/gpt-4.1 --case all --out /tmp/model-test.json

Inputs

  • --agent <name>: Target agent (chat, menial, coder, etc.)
  • --model <name>: Requested model alias/name to test
  • --case <id|all>: Case from references/test-cases.json or all
  • --timeout <sec>: Per-case timeout (default 120)
  • --out <file>: Optional JSON output file

Require at least one of --agent or --model.

What the runner does

  1. Load test cases from references/test-cases.json.
  2. Start openclaw logs --follow --json in parallel.
  3. Run openclaw agent --json with a bounded test prompt (asks agent to use a subagent for the task).
  4. Parse response + tailed logs.
  5. Emit machine-readable JSON and a short human summary.

Output format

Top-level JSON:

  • tool
  • timestamp
  • agent
  • requested_model
  • results[]

Each result entry returns:

  • test_case
  • agent
  • requested_model
  • actual_model (parsed from logs when available)
  • status (ok/error)
  • result_summary
  • runtime_seconds
  • tokens (when discoverable)
  • errors[]

Privacy & Safety

The tester spawns isolated subagent tasks with predefined test prompts only — no user data is passed to models. It tails OpenClaw logs to extract:

  • which model was actually selected (routing validation)
  • token usage statistics
  • runtime metrics

Log extraction uses regex patterns to find model/token fields. No personally identifiable information or arbitrary log content is captured — only structured fields related to the test execution.

Notes

  • Model extraction and token extraction are best-effort because log fields may vary by OpenClaw/provider version.
  • If openclaw config is invalid or gateway is unavailable, the script returns status=error with stderr details.
  • Edit references/test-cases.json to add custom prompts for your benchmark set.
  • All test cases are generic; no workspace or user data is baked in.

适合场景

01

OpenClaw 用户查找和安装 Skill 时

02

用户想查找某类 Agent Skill 时

03

需要根据任务场景推荐可安装能力包时

04

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

能力 5

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

OpenClaw

80.5%
按下载量换算3,124

安全审计

VirusTotal

通过

ClawScan

可疑

Static analysis

通过

权限和风险

只读

该 Skill 主要提供规则、说明或参考内容,本身偏只读;真正读写文件、联网或执行命令仍取决于宿主 Agent 的任务。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。来源安全扫描存在 warning/failed 结果,不能写成本站确认安全。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills