Token导航 LogoToken导航TokenDH.com
研究检索需要联网github未标认证来源可访问clear审计提醒

mteb-leaderboardmteb 排行榜

Agent Skill

mteb-leaderboard 用于查找、检索和筛选相关信息,适合在 Codex、Claude、Cursor、Gemini CLI 中需要根据关键词、任务场景或来源线索快速定位候选结果时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

791

周安装

32

GitHub Stars

93

下载量

248
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

3

许可证

MIT

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:mteb-leaderboard(mteb 排行榜)
来源仓库:https://github.com/letta-ai/skills
仓库路径:skills/mteb-leaderboard
安装命令:
npx skills add https://github.com/letta-ai/skills --skill mteb-leaderboard
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。不同来源提供的安装方式可能略有差异;本站展示可直接复制的安装命令,安装前请核对来源页面。

skills.shnpx skills
npx skills add https://github.com/letta-ai/skills --skill mteb-leaderboard

简介

mteb-leaderboard 用于查找、检索和筛选相关信息,适合快速定位候选结果。

  • 适用于需要根据关键词或任务场景从来源线索中获取信息的场景。
  • 通过 npx skills add 命令安装,需结合原始 README 核验具体用法。
  • 安装前建议确认权限范围、维护状态及是否触发联网或文件读写操作。
  • 适用宿主包括 Codex、Claude、Cursor、Gemini CLI,接入前应确认版本、权限和运行环境要求。

SKILL.md

MTEB Leaderboard Query Skill

This skill provides guidance for accurately querying machine learning model leaderboards and benchmarks, particularly the Massive Text Embedding Benchmark (MTEB) and related embedding leaderboards.

When to Use This Skill

  • Finding top-performing models on specific benchmarks (MTEB, Scandinavian Embedding Benchmark, etc.)
  • Answering questions about current leaderboard standings
  • Comparing model performance across different benchmarks
  • Tasks with specific temporal requirements (e.g., "as of August 2025")

Core Approach

Step 1: Identify Authoritative Data Sources

Before searching for results, establish which sources contain authoritative, current data:

  1. Primary Sources (prefer these):

- Official leaderboard websites (e.g., mteb-leaderboard on HuggingFace Spaces) - GitHub repositories with raw benchmark data - API endpoints or JSON data files from leaderboard maintainers

  1. Secondary Sources (use with caution):

- Academic papers (often outdated by publication time) - Blog posts and articles (may reference outdated results) - News articles about benchmark results

Step 2: Verify Temporal Alignment

When a task specifies a time constraint (e.g., "as of August 2025"):

  1. Check source publication/update dates - Academic papers are typically 6-18 months behind current leaderboard state
  2. Look for "last updated" timestamps on leaderboard pages
  3. Never assume paper results reflect current standings without verification
  4. Be explicit about temporal gaps - If using data from June 2024 to answer about August 2025, this is a 14+ month gap that likely invalidates the data

Step 3: Access Live Leaderboard Data

When web pages don't render properly (interactive charts, JavaScript-heavy pages):

  1. Look for raw data endpoints:

- Check for /api/ or /data/ endpoints - Search for JSON files in the page source - Look for GitHub repositories backing the leaderboard

  1. Try alternative access methods:

- HuggingFace Spaces often have Gradio APIs - Many leaderboards publish CSV/JSON exports - Check GitHub issues/discussions for data access tips

  1. Search for data repositories:

- site:github.com [leaderboard name] results json - site:huggingface.co [benchmark name] leaderboard

Step 4: Validate Model Eligibility

Do not make assumptions about which models "count" on a leaderboard:

  1. Check official leaderboard criteria - Some include API models, some don't
  2. Verify the answer format requirements against actual leaderboard entries
  3. Do not exclude models based on assumptions about what can be represented in a given format
  4. Consider all model types: open-source, API-based, fine-tuned variants

Verification Strategies

Cross-Reference Multiple Sources

  • Compare results from at least 2-3 independent sources
  • If sources disagree, prioritize the most recent authoritative source
  • Document discrepancies and their potential causes

Sanity Check Results

  • Verify the model actually appears on the leaderboard
  • Confirm the model name/organization format matches the source
  • Check if the model was released before the specified date

Test Alternative Access Methods

When primary access fails:

  1. Try the Wayback Machine for historical snapshots
  2. Search for leaderboard maintainer announcements
  3. Look for community discussions about recent changes
  4. Check if there's a programmatic API

Common Pitfalls to Avoid

1. Relying on Outdated Academic Papers

Academic papers have publication delays of 3-12 months. A paper published in June 2024 contains data from early 2024 at best. Never use paper results for questions about current standings.

2. Giving Up When Web Scraping Fails

Interactive leaderboards often don't render in simple web fetches. Always try:

  • Looking for underlying data files
  • Checking GitHub repositories
  • Finding API endpoints
  • Searching for data exports

3. Making Assumptions About Model Format

Do not assume API models (OpenAI, Cohere, etc.) cannot be valid answers. Check the actual task requirements and leaderboard contents.

4. Premature Conclusion Without Verification

Before writing a final answer:

  • Verify the model appears on the actual leaderboard
  • Confirm the ranking is current
  • Check that the model meets all task requirements

5. Ignoring Temporal Requirements

If a task asks about a specific date, ensure data sources reflect that timeframe. A 14-month gap between data and required date is unacceptable.

Systematic Search Strategy

When searching for leaderboard information:

  1. Start broad, then narrow:

- [benchmark name] leaderboard 2025 - [benchmark name] top models current - site:huggingface.co [benchmark name]

  1. Search for raw data:

- [benchmark name] results github - [benchmark name] json data - [benchmark name] api

  1. Search for recent updates:

- [benchmark name] new top model [current year] - [benchmark name] leaderboard update

  1. Avoid repetitive similar queries - If a query pattern isn't working after 2-3 attempts, change the approach rather than making minor variations

Output Checklist

Before submitting an answer, verify:

  • Data source is current (not outdated paper)
  • Model appears on the actual leaderboard
  • Temporal requirements are met
  • Model format matches requirements
  • No unvalidated assumptions were made
  • Answer was cross-referenced where possible

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

04

需要参考平台分布和安装热度时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

能力 5

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Claude Code

27.19%
按下载量换算67

Gemini CLI

23.64%
按下载量换算59

Antigravity

17.07%
按下载量换算42

windsurf

13.4%
按下载量换算33

OpenCode

7.98%
按下载量换算20

Codex

3.21%
按下载量换算8

安全审计

Gen Agent Trust Hub

通过

Socket

通过

Snyk

可疑

权限和风险

需要联网

该 Skill 可能需要联网访问来源站点、仓库或外部 API;具体网络访问范围需要结合源码和 README 复核。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。来源安全扫描存在 warning/failed 结果,不能写成本站确认安全。

来源信息

继续浏览同类 Skills