Token导航 LogoToken导航TokenDH.com
研究检索需要联网github未标认证来源可访问clear审计未展示

llm-testingLLM 测试

Agent Skill

用于辅助测试设计、自动化测试、用例整理和回归验证。它适合让 Agent 编写单元测试、端到端测试、测试计划或根据失败日志定位问题。使用时需要确认项目测试框架、运行命令和夹具数据,避免为了通过测试而改坏真实逻辑;涉及浏览器或外部服务时,应区分本地模拟、测试环境和生产环境。

总安装

428

周安装

18

GitHub Stars

公开资料未说明

下载量

150
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

MIT

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:llm-testing(LLM 测试)
来源仓库:https://github.com/yonatangross/skillforge-claude-plugin
仓库路径:skills/llm-testing
安装命令:
npx skills add yonatangross/skillforge-claude-plugin --skill "llm-testing"
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

AgentSkills.tonpx skills
npx skills add yonatangross/skillforge-claude-plugin --skill "llm-testing"

简介

发现并安装 AI 代理的技能。

  • 适用于 LLM 测试用例设计和验证场景。
  • 支持多种宿主环境的技能集成。
  • 安装命令:npx skills add yonatangross/skillforge-claude-plugin --skill "llm-testing"。
  • 注意确认权限范围及是否触发测试执行操作。

SKILL.md

LLM Testing Patterns

Test AI applications with deterministic patterns using DeepEval and RAGAS.

Quick Reference

Mock LLM Responses

from unittest.mock import AsyncMock, patch

@pytest.fixture
def mock_llm():
    mock = AsyncMock()
    mock.return_value = {"content": "Mocked response", "confidence": 0.85}
    return mock

@pytest.mark.asyncio
async def test_with_mocked_llm(mock_llm):
    with patch("app.core.model_factory.get_model", return_value=mock_llm):
        result = await synthesize_findings(sample_findings)
    assert result["summary"] is not None

DeepEval Quality Testing

from deepeval import assert_test
from deepeval.test_case import LLMTestCase
from deepeval.metrics import AnswerRelevancyMetric, FaithfulnessMetric

test_case = LLMTestCase(
    input="What is the capital of France?",
    actual_output="The capital of France is Paris.",
    retrieval_context=["Paris is the capital of France."],
)

metrics = [
    AnswerRelevancyMetric(threshold=0.7),
    FaithfulnessMetric(threshold=0.8),
]

assert_test(test_case, metrics)

Timeout Testing

import asyncio
import pytest

@pytest.mark.asyncio
async def test_respects_timeout():
    with pytest.raises(asyncio.TimeoutError):
        async with asyncio.timeout(0.1):
            await slow_llm_call()

Quality Metrics (2026)

MetricThresholdPurpose
Answer Relevancy≥ 0.7Response addresses question
Faithfulness≥ 0.8Output matches context
Hallucination≤ 0.3No fabricated facts
Context Precision≥ 0.7Retrieved contexts relevant

Anti-Patterns (FORBIDDEN)

# ❌ NEVER test against live LLM APIs in CI
response = await openai.chat.completions.create(...)

# ❌ NEVER use random seeds (non-deterministic)
model.generate(seed=random.randint(0, 100))

# ❌ NEVER skip timeout handling
await llm_call()  # No timeout!

# ✅ ALWAYS mock LLM in unit tests
with patch("app.llm", mock_llm):
    result = await function_under_test()

# ✅ ALWAYS use VCR.py for integration tests
@pytest.mark.vcr()
async def test_llm_integration():
    ...

Key Decisions

DecisionRecommendation
Mock vs VCRVCR for integration, mock for unit
TimeoutAlways test with < 1s timeout
Schema validationTest both valid and invalid
Edge casesTest all null/empty paths
Quality metricsUse multiple dimensions (3-5)

Detailed Documentation

ResourceDescription
references/deepeval-ragas-api.mdDeepEval & RAGAS API reference
examples/test-patterns.mdComplete test examples
checklists/llm-test-checklist.mdSetup and review checklists
scripts/llm-test-template.pyStarter test template

Related Skills

  • vcr-http-recording - Record LLM responses
  • llm-evaluation - Quality assessment
  • unit-testing - Test fundamentals

Capability Details

llm-response-mocking

Keywords: mock LLM, fake response, stub LLM, mock AI Solves:

  • Mock LLM responses in tests
  • Create deterministic AI test fixtures
  • Avoid live API calls in CI

async-timeout-testing

Keywords: timeout, async test, wait for, polling Solves:

  • Test async LLM operations
  • Handle timeout scenarios
  • Implement polling assertions

structured-output-validation

Keywords: structured output, JSON validation, schema validation, output format Solves:

  • Validate structured LLM output
  • Test JSON schema compliance
  • Assert output structure

deepeval-assertions

Keywords: DeepEval, assert_test, LLMTestCase, metric assertion Solves:

  • Use DeepEval for LLM assertions
  • Implement metric-based tests
  • Configure quality thresholds

golden-dataset-testing

Keywords: golden dataset, golden test, reference output, expected output Solves:

  • Test against golden datasets
  • Compare with reference outputs
  • Implement regression testing

vcr-recording

Keywords: VCR, cassette, record, replay, HTTP recording Solves:

  • Record LLM API responses
  • Replay recordings in tests
  • Create deterministic test suites

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

04

需要参考平台分布和安装热度时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Claude Code

26.44%
按下载量换算40

OpenCode

21.96%
按下载量换算33

Antigravity

16.65%
按下载量换算25

Gemini CLI

13.01%
按下载量换算20

windsurf

8.42%
按下载量换算13

trae

3.86%
按下载量换算6

安全审计

暂无安全审计结果可展示。

权限和风险

需要联网

该 Skill 可能需要联网访问来源站点、仓库或外部 API;具体网络访问范围需要结合源码和 README 复核。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills