Token导航 LogoToken导航TokenDH.com
研究检索需要联网clawhub未标认证来源可访问clear审计通过

robots-txt机器人.txt

Agent Skill

robots-txt 用于处理浏览器自动化、网页检查和页面信息提取,适合在 OpenClaw 中需要让 Agent 打开页面、读取网页或验证前端流程时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

2,717

周安装

111

GitHub Stars

公开资料未说明

下载量

870
OpenClaw

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

MIT-0

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:robots-txt(机器人.txt)
来源仓库:https://github.com/kostja94/robots-txt
安装命令:
openclaw skills install robots-txt
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 OpenClaw 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

ClawHubOpenClaw
openclaw skills install robots-txt

简介

robots-txt 用于解析和管理网站的 robots.txt 文件,支持爬虫规则验证与前端流程测试。

  • 适用于网页抓取、SEO 审核或浏览器自动化测试等需要遵守站点访问规则的场景。
  • 可读取现有规则并生成合规建议,辅助判断哪些路径可被安全访问。
  • 操作前应确认目标站点的访问权限,避免触发反爬机制或法律风险。
  • 建议结合真实页面访问测试结果综合评估规则影响。

SKILL.md

name
robots-txt
description
When the user wants to configure, audit, or optimize robots.txt. Also use when the user mentions "robots.txt," "crawler rules," "block crawlers," "AI crawlers," "GPTBot," "allow/disallow," "disallow path," "crawl directives," "user-agent," "block Googlebot," "fix robots.txt," "robots.txt blocking," or "search engine crawling." For indexing, use indexing.
metadata
version
1.1.1

SEO Technical: robots.txt

Guides configuration and auditing of robots.txt for search engine and AI crawler control.

When invoking: On first use, if helpful, open with 1–2 sentences on what this skill covers and why it matters, then provide the main output. On subsequent use or when the user asks to skip, go directly to the main output.

Scope (Technical SEO)

  • Robots.txt: Configure Disallow/Allow, Sitemap, Clean-param; audit for accidental blocks
  • Crawler access: Path-level crawl control; AI crawler allow/block strategy
  • Differentiation: robots.txt = crawl control (who accesses what paths); noindex = index control (what gets indexed). See indexing for page-level exclusions.

Initial Assessment

Check for project context first: If .claude/project-context.md or .cursor/project-context.md exists, read it for site URL and indexing goals.

Identify:

  1. Site URL: Base domain (e.g., https://example.com)
  2. Indexing scope: Full site, partial, or specific paths to exclude
  3. AI crawler strategy: Allow search/indexing vs. block training data crawlers

Best Practices

Purpose and Limitations

PointNote
PurposeControls crawler access; does NOT prevent indexing (disallowed URLs may still appear in search without snippet)
AdvisoryRules are advisory; malicious crawlers may ignore
Publicrobots.txt is publicly readable; use noindex or auth for sensitive content. See indexing

Crawl vs Index vs Link Equity (Quick Reference)

ToolControlsPrevents indexing?
robots.txtCrawl (path-level)No—blocked URLs may still appear in SERP
noindex (meta / X-Robots-Tag)Index (page-level)Yes. See indexing
nofollowLink equity onlyNo—does not control indexing

When to Use robots.txt vs noindex

UseToolExample
Path-level (whole directory)robots.txtDisallow: /admin/, Disallow: /api/, Disallow: /staging/
Page-level (specific pages)noindex meta / X-Robots-TagLogin, signup, thank-you, 404, legal. See indexing for full list
CriticalDo NOT block in robots.txtPages that use noindex—crawlers must access the page to read the directive

Paths to block in robots.txt: /admin/, /api/, /staging/, temp files. Paths to use noindex (allow crawl): /login/, /signup/, /thank-you/, etc.—see indexing.

Location and Format

ItemRequirement
PathSite root: https://example.com/robots.txt
EncodingUTF-8 plain text
StandardRFC 9309 (Robots Exclusion Protocol)

Core Directives

DirectivePurposeExample
User-agent:Target crawlerUser-agent: Googlebot, User-agent: *
Disallow:Block path prefixDisallow: /admin/
Allow:Allow path (can override Disallow)Allow: /public/
Sitemap:Declare sitemap absolute URLSitemap: https://example.com/sitemap.xml
Clean-param:Strip query params (Yandex)See below

Critical: Do Not Block

Do not blockReason
CSS, JS, imagesGoogle needs them to render pages; blocking breaks indexing
/_next/ (Next.js)Breaks CSS/JS loading; static assets in GSC "Crawled - not indexed" is expected. See indexing
Pages that use noindexCrawlers must access the page to read the noindex directive; blocking in robots.txt prevents that

Only block: paths that don't need crawling: /admin/, /api/, /staging/, temp files.

AI Crawler Strategy

robots.txt is effective for all measured AI crawlers (Vercel/MERJ study, 2024). Set rules per user-agent; check each vendor's docs for current tokens.

User-agentPurposeTypical
OAI-SearchBotChatGPT searchAllow
GPTBotOpenAI trainingDisallow
Claude-SearchBotClaude searchAllow
ClaudeBotAnthropic trainingDisallow
PerplexityBotPerplexity searchAllow
Google-ExtendedGemini trainingDisallow
CCBotCommon Crawl (LLM training)Disallow
BytespiderByteDanceDisallow
Meta-ExternalAgentMetaDisallow
AppleBotApple (Siri, Spotlight); renders JSAllow for indexing

Allow vs Disallow: Allow search/indexing bots (OAI-SearchBot, Claude-SearchBot, PerplexityBot); Disallow training-only bots (GPTBot, ClaudeBot, CCBot) if you don't want content used for model training. See site-crawlability for AI crawler optimization (SSR, URL management).

Clean-param (Yandex)

Clean-param: utm_source&utm_medium&utm_campaign&utm_term&utm_content&ref&fbclid&gclid

Output Format

  • Current state (if auditing)
  • Recommended robots.txt (full file)
  • Compliance checklist
  • References: Google robots.txt

Related Skills

  • indexing: Full noindex page-type list; when to use noindex vs robots.txt; GSC indexing diagnosis
  • page-metadata: Meta robots (noindex, nofollow) implementation
  • xml-sitemap: Sitemap URL to reference in robots.txt
  • site-crawlability: Broader crawl and structure guidance; AI crawler optimization
  • rendering-strategies: SSR, SSG, CSR; content in initial HTML for crawlers

适合场景

01

OpenClaw 用户查找和安装 Skill 时

02

用户想查找某类 Agent Skill 时

03

需要根据任务场景推荐可安装能力包时

04

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

能力 5

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

OpenClaw

71.77%
按下载量换算624

安全审计

VirusTotal

通过

ClawScan

通过

Static analysis

通过

权限和风险

需要联网

该 Skill 可能需要联网访问来源站点、仓库或外部 API;具体网络访问范围需要结合源码和 README 复核。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills