Token导航 LogoToken导航TokenDH.com
开发敏感数据github未标认证来源可访问许可证需确认审计提醒

langsmith-evaluator朗史密斯评估器

Agent Skill

langsmith-evaluator 用于处理 GitHub 仓库、Issue、Pull Request 和代码协作信息,适合在 Codex、Claude、Cursor、Gemini CLI 中需要围绕仓库状态、代码变更或协作事项进行整理时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

34,920

周安装

1,508

GitHub Stars

103

下载量

12,240
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:langsmith-evaluator(朗史密斯评估器)
来源仓库:https://github.com/langchain-ai/langsmith-skills
仓库路径:skills/langsmith-evaluator
安装命令:
npx skills add https://github.com/langchain-ai/langsmith-skills --skill langsmith-evaluator
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

skills.shnpx skills
npx skills add https://github.com/langchain-ai/langsmith-skills --skill langsmith-evaluator

简介

使用法学硕士作为法官和自定义代码评估器为 LangSmith 构建评估管道。

  • 三个核心组件:创建评估器(LLM-as-Judge 或自定义代码)、定义运行函数以捕获代理输出和轨迹、在本地运行评估或通过上传的评估器自动运行
  • 支持离线评估器(将运行输出与数据集示例进行比较)和在线评估器(对生产运行进行实时质量检查)
  • 需要 LangSmith API 密钥和项目配置;包括 Python 和 TypeScript 示例,为 LLM 法官提供结构化输出支持
  • 关键工作流程:在编写评估器之前检查实际的代理输出结构和数据集模式;查询 LangSmith 轨迹以验证轨迹数据和字段名称是否匹配

SKILL.md

LANGSMITH_API_KEY=lsv2_pt_your_api_key_here          # REQUIRED
LANGSMITH_PROJECT=your-project-name                   # Check this to know which project has traces
LANGSMITH_WORKSPACE_ID=your-workspace-id              # Optional: for org-scoped keys
OPENAI_API_KEY=your_openai_key                        # For LLM as Judge

Authentication is REQUIRED: either set the LANGSMITH_API_KEY environment variable, or pass the --api-key flag to CLI commands (preferred):

langsmith evaluator list --api-key $LANGSMITH_API_KEY

IMPORTANT: Always check the environment variables or .env file for LANGSMITH_PROJECT before querying or interacting with LangSmith. This tells you which project contains the relevant traces and data. If the LangSmith project is not available, use your best judgement to identify the right one.

Python Dependencies

pip install langsmith langchain-openai python-dotenv

CLI Tool (for uploading evaluators)

curl -sSL https://raw.githubusercontent.com/langchain-ai/langsmith-cli/main/scripts/install.sh | sh

JavaScript Dependencies

npm install langsmith openai

<crucial_requirement>

Golden Rule: Inspect Before You Implement

CRITICAL: Before writing ANY evaluator or extraction logic, you MUST:

  1. Run your agent on sample inputs and capture the actual output
  2. Inspect the output - print it, query LangSmith traces, understand the exact structure
  3. Only then write code that processes that output

Output structures vary significantly by framework, agent type, and configuration. Never assume the shape - always verify first. Query LangSmith traces to when outputs don't contain needed data to understand how to extract from execution. </crucial_requirement>

<evaluator_format>

Offline vs Online Evaluators

Offline Evaluators (attached to datasets):

  • Function signature: (run, example) - receives both run outputs and dataset example
  • Use case: Comparing agent outputs to expected values in a dataset
  • Upload with: --dataset "Dataset Name"

Online Evaluators (attached to projects):

  • Function signature: (run) - receives only run outputs, NO example parameter
  • Use case: Real-time quality checks on production runs (no reference data)
  • Upload with: --project "Project Name"

CRITICAL - Return Format:

  • Each evaluator returns ONE metric only. For multiple metrics, create multiple evaluator functions.
  • Do NOT return {"metric_name": value} or lists of metrics - this will error.

CRITICAL - Local vs Uploaded Differences:

Local evaluate()Uploaded to LangSmith
Column namePython: auto-derived from function name. TypeScript: must include key field or column is untitledComes from evaluator name set at upload time. Do NOT include key — it creates a duplicate column
Python run typeRunTree object → run.outputs (attribute)dictrun["outputs"] (subscript). Handle both: run.outputs if hasattr(run, "outputs") else run.get("outputs", {})
TypeScript run typeAlways attribute access: run.outputs?.fieldAlways attribute access: run.outputs?.field
Python return{"score": value, "comment": "..."}{"score": value, "comment": "..."}
TypeScript return{key: "name", score: value, comment: "..."}{score: value, comment: "..."}
</evaluator_format>

<evaluator_types>

  • LLM as Judge - Uses an LLM to grade outputs. Best for subjective quality (accuracy, helpfulness, relevance).
  • Custom Code - Deterministic logic. Best for objective checks (exact match, trajectory validation, format compliance). </evaluator_types>

<llm_judge>

LLM as Judge Evaluators

NOTE: LLM-as-Judge upload is currently not supported by the CLI — only code evaluators are supported. For evaluations against a dataset, STRONGLY PREFER defining local evaluators to use with evaluate(evaluators=[...]).

class Grade(TypedDict): reasoning: Annotated[str,..., "Explain your reasoning"] is_accurate: Annotated[bool,..., "True if response is accurate"]

judge = ChatOpenAI(model="gpt-4o-mini", temperature=0).with_structured_output(Grade, method="json_schema", strict=True)

async def accuracy_evaluator(run, example): run_outputs = run.outputs if hasattr(run, "outputs") else run.get("outputs", {}) or {} example_outputs = example.outputs if hasattr(example, "outputs") else example.get("outputs", {}) or {} grade = await judge.ainvoke([{"role": "user", "content": f"Expected: {example_outputs}\nActual: {run_outputs}\nIs this accurate?"}]) return {"score": 1 if grade["is_accurate"] else 0, "comment": grade["reasoning"]}

</python>

<typescript>

import OpenAI from "openai";

const openai = new OpenAI();

async function accuracyEvaluator(run, example) { const runOutputs = run.outputs ?? {}; const exampleOutputs = example.outputs ?? {};

const response = await openai.chat.completions.create({ model: "gpt-4o-mini", temperature: 0, response_format: { type: "json_object" }, messages: [ { role: "system", content: 'Respond with JSON: {"is_accurate": boolean, "reasoning": string}' }, { role: "user", content: Expected: ${JSON.stringify(exampleOutputs)}\nActual: ${JSON.stringify(runOutputs)}\nIs this accurate? } ] });

const grade = JSON.parse(response.choices[0].message.content); return { score: grade.is_accurate ? 1 : 0, comment: grade.reasoning }; }


<code_evaluators>

## Custom Code Evaluators

**Before writing an evaluator:**

1. Inspect your dataset to understand expected field names (see Golden Rule above)
2. Test your run function and verify its output structure matches the dataset schema
3. Query LangSmith traces to debug any mismatches

<run_functions>

## Defining Run Functions

Run functions execute your agent and return outputs for evaluation.

**CRITICAL - Test Your Run Function First:** Before writing evaluators, you MUST test your run function and inspect the actual output structure. Output shapes vary by framework, agent type, and configuration.

**Debugging workflow:**

1. Run your agent once on sample input
2. Query the trace to see the execution structure
3. Print the raw output and verify against trace to output contains the right data
4. Adjust the run function as needed
5. Verify your output matches your dataset schema

**Try your hardest to match your run function output to your dataset schema.** This makes evaluators simple and reusable. If matching isn't possible, your evaluator must know how to extract and compare the right fields from each side.

### Capturing Trajectories

For trajectory evaluation, your run function must capture tool calls during execution.

**CRITICAL:** Run output formats vary significantly by framework and agent type. You MUST inspect before implementing:

**LangGraph agents (LangChain OSS):** Use `stream_mode="debug"` with `subgraphs=True` to capture nested subagent tool calls.

import uuid

def run_agent_with_trajectory(agent, inputs: dict) -> dict: config = {"configurable": {"thread_id": f"eval-{uuid.uuid4()}"}} trajectory = [] final_result = None

for chunk in agent.stream(inputs, config=config, stream_mode="debug", subgraphs=True): # STEP 1: Print chunks to understand the structure print(f"DEBUG chunk: {chunk}")

# STEP 2: Write extraction based on YOUR observed structure # ... your extraction logic here ...

# IMPORTANT: After running, query the LangSmith trace to verify # your trajectory data is complete. Default output may be missing # tool calls that appear in the trace. return {"output": final_result, "trajectory": trajectory}


**Custom / Non-LangChain Agents:**

1. **Inspect output first** - Run your agent and inspect the result structure. Trajectory data may already be included in the output (e.g., `result.tool_calls`, `result.steps`, etc.)
2. **Callbacks/Hooks** - If your framework supports execution callbacks, register a hook that records tool names on each invocation
3. **Parse execution logs** - As a last resort, extract tool names from structured logs or trace data

The key is to capture the tool name at execution time, not at definition time. </run_functions>

**IMPORTANT - Auto-Run Behavior:** Evaluators uploaded to a dataset **automatically run** when you run experiments on that dataset. You do NOT need to pass them to `evaluate()` - just run your agent against the dataset and the uploaded evaluators execute automatically.

**IMPORTANT - Local vs Uploaded:** Uploaded evaluators run in a sandboxed environment with very limited package access. Only use built-in/standard library imports, and place all imports **inside** the evaluator function body. For dataset (offline) evaluators, prefer running locally with `evaluate(evaluators=[...])` first — this gives you full package access.

**IMPORTANT - Code vs Structured Evaluators:**

- **Code evaluators** (what the CLI uploads): Run in a limited environment without external packages. Use for deterministic logic (exact match, trajectory validation).
- **Structured evaluators** (LLM-as-Judge): Configured via LangSmith UI, use a specific payload format with model/prompt/schema. The CLI does not support this format yet.

**IMPORTANT - Choose the right target:**

- `--dataset`: Offline evaluator with `(run, example)` signature - for comparing to expected values
- `--project`: Online evaluator with `(run)` signature - for real-time quality checks

You must specify one. Global evaluators are not supported.

List all evaluators

langsmith evaluator list --api-key $LANGSMITH_API_KEY

Upload offline evaluator (attached to dataset)

langsmith evaluator upload my_evaluators.py \ --name "Trajectory Match" --function trajectory_evaluator \ --dataset "My Dataset" --replace --api-key $LANGSMITH_API_KEY

Upload online evaluator (attached to project)

langsmith evaluator upload my_evaluators.py \ --name "Quality Check" --function quality_check \ --project "Production Agent" --replace --api-key $LANGSMITH_API_KEY

Delete

langsmith evaluator delete "Trajectory Match" --api-key $LANGSMITH_API_KEY


**IMPORTANT - Safety Prompts:**

- The CLI prompts for confirmation before destructive operations
- **NEVER use `--yes` flag unless the user explicitly requests it**

<best_practices>

1. **Use structured output for LLM judges** - More reliable than parsing free-text
2. **Match evaluator to dataset type**
  - Final Response → LLM as Judge for quality
  - Trajectory → Custom Code for sequence
3. **Use async for LLM judges** - Enables parallel evaluation
4. **Test evaluators independently** - Validate on known good/bad examples first
5. **Choose the right language**
  - Python: Use for Python agents, langchain integrations
  - JavaScript: Use for TypeScript/Node.js agents </best_practices>

<running_evaluations>

## Running Evaluations

**Uploaded evaluators** auto-run when you run experiments - no code needed. **Local evaluators** are passed directly for development/testing.

# Uploaded evaluators run automatically

results = evaluate(run_agent, data="My Dataset", experiment_prefix="eval-v1")

# Or pass local evaluators for testing

results = evaluate(run_agent, data="My Dataset", evaluators=[my_evaluator], experiment_prefix="eval-v1")

</python>

<typescript>

import { evaluate } from "langsmith/evaluation";

// Uploaded evaluators run automatically
const results = await evaluate(runAgent, {
  data: "My Dataset",
  experimentPrefix: "eval-v1",
});

// Or pass local evaluators for testing
const results = await evaluate(runAgent, {
  data: "My Dataset",
  evaluators: [myEvaluator],
  experimentPrefix: "eval-v1",
});

Output doesn't match what you expect: Query the LangSmith trace. It shows exact inputs/outputs at each step - compare what you find to what you're trying to extract.

One metric per evaluator: Return {"score": value, "comment": "..."}. For multiple metrics, create separate functions.

Field name mismatch: Your run function output must match dataset schema exactly. Inspect dataset first with client.read_example(example_id).

RunTree vs dict (Python only): Local evaluate() passes RunTree, uploaded evaluators receive dict. Handle both:

run_outputs = run.outputs if hasattr(run, "outputs") else run.get("outputs", {}) or {}

TypeScript always uses attribute access: run.outputs?.field

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Codex

37.4%
按下载量换算4,578

Claude

26.75%
按下载量换算3,274

Cursor

18.98%
按下载量换算2,323

Gemini CLI

9.73%
按下载量换算1,191

安全审计

Gen Agent Trust Hub

通过

Socket

通过

Snyk

可疑

权限和风险

敏感数据

该 Skill 可能接触密钥、Token、环境变量或敏感配置,应进入高风险复核队列,默认不自动发布。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。来源安全扫描存在 warning/failed 结果,不能写成本站确认安全。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills