Token导航 LogoToken导航TokenDH.com
研究检索敏感数据github未标认证来源可访问许可证需确认审计通过

ai-bot-log-auditai 机器人日志审计

Agent Skill

用于辅助安全审计、权限检查、凭据风险、认证流程和常见漏洞排查。它适合让 Agent 梳理敏感配置、检查依赖风险、分析鉴权逻辑或生成安全复核清单。使用时不能把工具输出直接当最终结论,涉及密钥、令牌、用户数据或生产系统时,应先确认最小权限、脱敏方式和操作边界。

总安装

1,123

周安装

45

GitHub Stars

85

下载量

364
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:ai-bot-log-audit(ai 机器人日志审计)
来源仓库:https://github.com/guia-matthieu/clawfu-skills
仓库路径:skills/ai-bot-log-audit
安装命令:
npx skills add https://github.com/guia-matthieu/clawfu-skills --skill ai-bot-log-audit
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

skills.shnpx skills
npx skills add https://github.com/guia-matthieu/clawfu-skills --skill ai-bot-log-audit

简介

用于辅助安全审计、权限检查、凭据风险、认证流程和常见漏洞排查。

  • 适合让 Agent 梳理敏感配置、检查依赖风险、分析鉴权逻辑或生成安全复核清单。
  • 使用时不能把工具输出直接当最终结论。
  • 涉及密钥、令牌、用户数据或生产系统时,应先确认最小权限、脱敏方式和操作边界。
  • ai-bot-log-audit 属于研究检索类 Skill,可作为该场景下的辅助能力补充。

SKILL.md

AI Bot Log Audit

Analyze server logs to understand how AI crawlers retrieve your content, then optimize placement and structure for maximum citation probability. Based on Metehan Yeşilyurt's log file analysis framework.

When to Use This Skill

Use this skill when you need to:

  • Audit AI bot crawl patterns on your site (what they fetch, how often, what they skip)
  • Diagnose citation gaps — your content exists but AI search doesn't cite it
  • Optimize content placement for LLM retrieval mechanics (lost-in-the-middle, embedding similarity)
  • Compare AI bot behavior across different products (Google, OpenAI, Anthropic, Perplexity)
  • Build a GEO strategy grounded in actual crawl data, not assumptions
  • Identify which pages AI bots prioritize and which they ignore

This skill is particularly valuable for:

  • SEO professionals adding AI search to their optimization scope
  • Publishers monitoring AI traffic and content extraction
  • Technical SEOs conducting log file analysis
  • Content strategists deciding what to optimize for AI visibility

Methodology Foundation

Source Expert: Metehan Yeşilyurt (SEO Consultant, speaker at BrightonSEO, The Search Session)

Core Thesis: Log file analysis reveals the real mechanics of AI search retrieval. Different AI products (AI Overviews, AI Mode, Perplexity) use fundamentally different retrieval pipelines, so optimizing for "AI search" requires understanding each product's specific crawl and retrieval behavior.

"AI Overviews, AI Mode, and Web Guide are three completely different products with different retrieval behaviors. You can't optimize for 'AI search' as if it's one thing." — Metehan Yeşilyurt, The Search Session

Key Insight: The shift from traditional SEO to AI search optimization is from chasing clicks in deterministic SERPs to chasing citations in machine-generated text. Log files reveal the mechanics behind citation selection.


What Claude Does vs What You Decide

Claude DoesYou Decide
Guides log analysis methodologyWhich log data to provide
Identifies patterns in crawl behaviorStrategic priority of AI products
Recommends content placement optimizationsContent creation/modification scope
Maps retrieval behavior to optimization actionsResource allocation for changes
Creates audit templates and checklistsWhich recommendations to implement

What This Skill Does

When invoked, I will guide you through:

  1. Log Collection — Identify and extract AI bot activity from server logs
  2. Bot Identification — Map user agents to AI products
  3. Crawl Pattern Analysis — Understand what AI bots fetch and ignore
  4. Retrieval Mechanics — Learn how each AI product processes your content
  5. Optimization Actions — Specific changes to improve AI citation probability

Instructions

Phase 1: AI Bot Identification

Known AI Crawlers (2026)

Bot User AgentOperatorPurposeRespects robots.txt?
GPTBotOpenAITraining data + ChatGPT BrowseYes
ChatGPT-UserOpenAIReal-time browsing in ChatGPTYes
OAI-SearchBotOpenAISearchGPT / ChatGPT SearchYes
ClaudeBotAnthropicTraining data collectionYes
PerplexityBotPerplexityReal-time search + answer generationYes (mostly)
Google-ExtendedGoogleGemini/AI training (deprecated — now part of Googlebot)N/A
GooglebotGoogleTraditional crawl + AI Overviews + AI ModeYes
BytespiderByteDanceTraining for TikTok AI featuresYes
CCBotCommon CrawlOpen dataset used by many AI modelsYes
Applebot-ExtendedAppleApple Intelligence featuresYes

Log Extraction Commands

# Extract all AI bot hits from Apache/Nginx access logs
grep -iE "(GPTBot|ChatGPT-User|OAI-SearchBot|ClaudeBot|PerplexityBot|Bytespider|CCBot|Applebot-Extended)" access.log > ai_bots.log

# Count hits by bot
awk -F'"' '{print $6}' ai_bots.log | grep -oE "(GPTBot|ChatGPT-User|OAI-SearchBot|ClaudeBot|PerplexityBot|Bytespider|CCBot)" | sort | uniq -c | sort -rn

# Top pages crawled by AI bots
awk '{print $7}' ai_bots.log | sort | uniq -c | sort -rn | head -50

# Crawl frequency by day
awk '{print $4}' ai_bots.log | cut -d: -f1 | tr -d '[' | sort | uniq -c

# Response codes for AI bots (are they getting 200s or errors?)
awk '{print $9}' ai_bots.log | sort | uniq -c | sort -rn

Phase 2: Crawl Pattern Analysis

What to Look For

PatternWhat It MeansAction
Bot fetches page frequentlyPage is in retrieval index, content mattersOptimize this page first
Bot fetches page once then stopsPage was evaluated and deprioritizedImprove content quality/freshness
Bot never fetches a pagePage not discovered or blockedCheck internal linking, sitemap, robots.txt
Bot gets 404/500Technical issue blocking retrievalFix immediately
Bot fetches but doesn't citeContent retrieved but not selected for answersImprove structure, uniqueness, authority

Crawl Budget Analysis

AI bots have crawl budgets like traditional bots. If your site is large:

QUESTIONS TO ANSWER:
1. What % of pages are being crawled by AI bots?
2. Are high-value pages being fetched?
3. Are AI bots wasting time on low-value pages (tag pages, pagination)?
4. What's the crawl frequency — daily? weekly? sporadic?
5. Do different AI bots prioritize different pages?

Comparative Bot Behavior

Map how each bot interacts with your site differently:

ANALYSIS TEMPLATE:

GPTBot:
  Pages fetched: ___
  Top pages: ___
  Crawl frequency: ___
  Response codes: ___

PerplexityBot:
  Pages fetched: ___
  Top pages: ___
  Crawl frequency: ___
  Response codes: ___

ClaudeBot:
  Pages fetched: ___
  Top pages: ___
  Crawl frequency: ___
  Response codes: ___

Phase 3: Understanding Retrieval Mechanics

Each AI product uses a different retrieval pipeline. Optimizing requires understanding the mechanics.

Perplexity: Embedding Similarity + Source Diversification

Perplexity uses high embedding similarity to match content to queries, then actively diversifies sources.

What this means for you:

  • Semantic match matters more than keyword match
  • Being the single best source isn't enough — Perplexity diversifies
  • Your content needs to be retrievable AND complementary to other sources
  • Structured, extractable content wins (tables, definitions, clear claims)

Google AI Overviews: Search Index + LLM Layer

AI Overviews sit on top of existing Google Search rankings. Your traditional SEO signals matter.

What this means for you:

  • If you don't rank on page 1 for traditional search, you're unlikely to appear in AI Overviews
  • The AI layer selects from already-ranked pages and synthesizes
  • E-E-A-T signals are inherited from search ranking

Google AI Mode: Multi-Step Query Fan-Out

AI Mode decomposes complex questions into sub-queries, retrieves for each, then synthesizes.

What this means for you:

  • Content that answers sub-questions gets included
  • Comprehensive topical coverage across your site matters
  • Internal linking between related topics helps AI Mode stitch answers together

Phase 4: LLM Vulnerability Awareness

Understanding known LLM processing weaknesses helps you position content for maximum retrieval.

4 Key LLM Vulnerabilities

VulnerabilityDescriptionOptimization
Recency biasLLMs weight recent information higher, even when older info is more accurateDate your content clearly. Update regularly. Recent timestamps = retrieval advantage
Lost-in-the-middleLLMs process beginnings and endings of context better than middlesPlace key claims, definitions, and statistics at the START and END of pages/sections
Data poisoning susceptibilityLLMs can be influenced by coordinated content patternsNot actionable for you, but be aware: competitors can game AI answers
Prompt injectionLLMs can be manipulated via hidden instructionsNot actionable for SEO, but relevant for security audits

Lost-in-the-Middle: Practical Content Placement

This is the most actionable vulnerability for SEO:

PAGE STRUCTURE FOR LLM RETRIEVAL:

TOP (HIGH ATTENTION):
├── Key definition / direct answer
├── Critical statistics with context
└── Main claim / thesis

MIDDLE (LOW ATTENTION):
├── Supporting details
├── Tables and structured data (LLMs handle these well regardless of position)
├── Examples and elaboration
└── Supporting evidence

BOTTOM (HIGH ATTENTION):
├── Summary of key points
├── FAQ with direct answers
├── Conclusion with standalone claims
└── Updated date
"FAQs at the end of the page, tables in the middle — that's the lost-in-the-middle fix. LLMs handle structured data well regardless of position, but flowing text in the middle gets deprioritized." — Metehan Yeşilyurt

Phase 5: Audit Report Template

AI Bot Log Audit Report

# AI Bot Log Audit: [Site Name]
**Date:** [Date]
**Period analyzed:** [Date range]
**Log source:** [Apache/Nginx/CDN]

## Executive Summary
[2-3 sentence overview of findings]

## Bot Activity Overview

| Bot | Total Hits | Unique Pages | Avg Daily Hits | Status |
|-----|-----------|-------------|----------------|--------|
| GPTBot | ___ | ___ | ___ | Active/Inactive |
| PerplexityBot | ___ | ___ | ___ | Active/Inactive |
| ClaudeBot | ___ | ___ | ___ | Active/Inactive |
| ChatGPT-User | ___ | ___ | ___ | Active/Inactive |

## Top Pages by AI Bot Activity
1. [URL] — [hits] hits — [which bots]
2. [URL] — [hits] hits — [which bots]
3. [URL] — [hits] hits — [which bots]

## Pages Not Crawled (Expected High-Value)
- [URL] — Likely cause: [blocked/not linked/low authority]
- [URL] — Likely cause: [___]

## Technical Issues
- [N] pages returning 404 to AI bots
- [N] pages returning 500 to AI bots
- robots.txt blocking: [yes/no — which bots?]

## Content Optimization Recommendations

### Priority 1: Fix Technical Issues
- [ ] [Specific fix]

### Priority 2: Optimize High-Traffic AI Pages
- [ ] Apply lost-in-the-middle structure to [page]
- [ ] Add entity schema to [page]
- [ ] Update freshness signals on [page]

### Priority 3: Improve Discoverability
- [ ] Internal link to [uncrawled page] from [crawled hub]
- [ ] Add to sitemap: [page]
- [ ] Create content for [gap topic]

## robots.txt Recommendations
[Current AI bot directives and recommended changes]

Examples

Example 1: Publisher Losing AI Traffic

Context: News publisher noticed declining traffic from AI referrals. Log audit reveals:

Findings:

  • GPTBot hits dropped 60% after robots.txt change (accidentally blocked)
  • PerplexityBot only crawling homepage and top 5 articles
  • Most content pages have zero AI bot visits
  • AI bots hitting 404s on URL structure that changed 3 months ago

Actions:

  1. Fix robots.txt — allow GPTBot and PerplexityBot
  2. Add 301 redirects for old URL structure
  3. Improve internal linking from high-crawl pages to deep content
  4. Add article schema with datePublished and dateModified
  5. Restructure top articles with lost-in-the-middle optimization

Example 2: SaaS Blog Not Getting AI Citations

Context: SaaS company publishes weekly blog but never appears in ChatGPT or Perplexity answers.

Findings:

  • ClaudeBot and GPTBot crawl regularly (technical access is fine)
  • Content is generic ("10 Tips for..." format) — interchangeable with competitors
  • No original data, case studies, or unique insights
  • Author pages have no Person schema
  • Brand search volume: 100/month (competitors: 3,000+)

Actions:

  1. Content strategy shift: original research > generic guides
  2. Add Person schema for all authors with credentials
  3. Place unique data/claims at start and end of articles (lost-in-the-middle)
  4. Build entity authority for 2 key authors (podcast appearances, guest posts)
  5. Create comparison content where their product is one of the compared options

Checklists & Templates

Quick Audit Checklist

ACCESS CHECK
[ ] AI bots not blocked in robots.txt (check each bot separately)
[ ] AI bots getting 200 responses (not 403/404/500)
[ ] Sitemap accessible and up to date
[ ] Key pages present in sitemap

CRAWL PATTERNS
[ ] High-value pages are being crawled by AI bots
[ ] Crawl frequency is consistent (not declining)
[ ] AI bots not wasting budget on low-value pages
[ ] Multiple AI bots active (not just one)

CONTENT OPTIMIZATION
[ ] Key claims at beginning and end of pages (not buried in middle)
[ ] Tables and structured data for comparisons
[ ] FAQ sections at bottom of major pages
[ ] Content is updated with recent dates
[ ] Unique data or original insights present

ENTITY SIGNALS
[ ] Author/organization schema on relevant pages
[ ] Entity mentioned consistently across the site
[ ] Cross-platform entity presence established

Skill Boundaries

What This Skill Does Well

  • Guiding server log analysis for AI bot patterns
  • Explaining retrieval mechanics for different AI products
  • Recommending content placement optimizations
  • Creating structured audit reports

What This Skill Cannot Do

  • Access or parse your actual server logs
  • Query live AI search results
  • Guarantee citation in AI responses
  • Replace technical SEO expertise for complex log analysis

References

Primary Sources:

  • Metehan Yeşilyurt — "Log Files, AI Bots, and the Real Mechanics of AI Search" (The Search Session, Advanced Web Ranking)
  • Mark Williams-Cook — Signal decay thesis and query fan-out analysis
  • Lily Ray — Entity-first framework and cross-platform authority

Tools Referenced:

  • Screaming Frog Log File Analyzer
  • GoAccess (open-source log analyzer)
  • queryfan.com (query fan-out analysis)

Technical References:

  • OpenAI GPTBot documentation
  • Anthropic ClaudeBot documentation
  • Google Search Central — Googlebot and AI features

Related Skills

  • llm-optimized-content — Content optimization for AI search (GEO)
  • entity-seo-playbook — Building entity authority for AI visibility
  • lighthouse-audit — Technical SEO and performance audit
  • seo-content-writer — SEO content creation with E-E-A-T
  • schema-markup — Structured data implementation

Skill Metadata

name: ai-bot-log-audit
category: seo-tools
subcategory: geo
version: 1.0.0
author: GUIA
source_expert: Metehan Yeşilyurt (SEO Consultant, BrightonSEO speaker) — The Search Session podcast
difficulty: intermediate
mode: centaur
tags: [ai-bots, log-analysis, geo, ai-search, crawl-audit, perplexity, gptbot, claudebot, retrieval, lost-in-the-middle]
created: 2026-02-10
updated: 2026-02-10

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Codex

36.81%
按下载量换算134

Claude

30.73%
按下载量换算112

Cursor

20.17%
按下载量换算73

Gemini CLI

9.64%
按下载量换算35

安全审计

Gen Agent Trust Hub

通过

Socket

通过

Snyk

通过

权限和风险

敏感数据

该 Skill 可能接触密钥、Token、环境变量或敏感配置,应进入高风险复核队列,默认不自动发布。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills