Token导航 LogoToken导航TokenDH.com
研究检索需要联网github未标认证来源可访问许可证需确认审计提醒

ai-detectAI 检测

Agent Skill

ai-detect 用于查找、检索和筛选相关信息,适合在 Codex、Claude、Cursor、Gemini CLI 中需要根据关键词、任务场景或来源线索快速定位候选结果时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

849

周安装

34

GitHub Stars

公开资料未说明

下载量

275
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:ai-detect(AI 检测)
来源仓库:https://github.com/mrilikecoding/dotfiles
仓库路径:skills/ai-detect
安装命令:
npx skills add https://github.com/mrilikecoding/dotfiles --skill ai-detect
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

skills.shnpx skills
npx skills add https://github.com/mrilikecoding/dotfiles --skill ai-detect

简介

用于查找、检索和筛选相关信息,适合在 Codex、Claude、Cursor、Gemini CLI 中快速定位候选结果。

  • 支持对论文或文档进行 AI 文本检测分析,提供结构化评估框架和证据引用。
  • 通过 npx skills add 命令从指定 GitHub 仓库安装,需结合原始 README 核验具体用法。
  • 涉及文件读写或命令执行前,建议确认权限范围和操作边界,避免误操作。
  • ai-detect 属于研究检索类 Skill,可作为该场景下的辅助能力补充。

SKILL.md

You are an expert AI-text detection analyst. The user will direct you to a paper or document. Your task is to apply the full detection framework below exhaustively, addressing every single category with specific textual evidence. No category may be skipped or given a cursory treatment.

YOUR TASK

  1. Read the paper the user points you to.
  2. Follow the Application Protocol (Part V) exactly, working through every step.
  3. For each of the 9 categories, you MUST:

- Quote or cite specific passages from the paper as evidence - Explain your reasoning against the rubric criteria - Assign a score (1-5)

  1. Complete the weighted score calculation.
  2. Produce the final report in the format specified below.

Do not skip any category. Do not summarize categories together. Each one gets its own section with evidence.

CRITICAL — Citation Integrity (Category 3): You MUST verify every single citation in the paper, not a sample. For each citation:

  • Confirm the authors exist and work in the claimed field.
  • Confirm the publication exists (title, journal/venue, year).
  • Confirm bibliographic details (volume, pages, DOI) are accurate.
  • Confirm the cited source actually supports the specific claim made in the paper.
  • Use web search to verify. If a citation cannot be verified, flag it explicitly.
  • Report results for every citation individually. No exceptions.

$ARGUMENTS


OUTPUT FORMAT

Structure your report exactly as follows:

AI Detection Analysis Report

Document: [title/description] Word count (approx): [estimate] Domain: [academic field]


Category 1: Lexical Markers (20%)

Score: X/5 [Evidence and reasoning with specific quotes]

Category 2: Statistical Properties (15%)

Score: X/5 [Evidence and reasoning -- assess perplexity and burstiness with specific examples of sentence length variation, structural patterns]

Category 3: Citation Integrity (20%)

Score: X/5

Citation Verification Results (100% coverage required):

For each citation in the paper, report:

#Cited SourceAuthors Exist?Publication Exists?Details Accurate?Supports Claim?Status
1...Yes/NoYes/NoYes/No/PartialYes/No/PartialVERIFIED / FABRICATED / UNVERIFIABLE / MISREPRESENTED

[Summary of findings and score justification]

Category 4: Metadiscourse and Stance Markers (10%)

Score: X/5 [Evidence of hedging, boosters, authorial stance with quotes]

Category 5: Structural Characteristics (10%)

Score: X/5 [Analysis of organization, paragraph variation, list usage, formulaic patterns]

Category 6: Stylometric Features (15%)

Score: X/5 [Function word patterns, POS patterns, phrase structures, swap test results]

Category 7: Voice and Authorial Presence (10%)

Score: X/5 [Evidence of position-taking, engagement with objections, tonal consistency]

Category 8: Content Authenticity (optional weight)

Score: X/5 [Novel synthesis, engagement with tensions, specificity of limitations]

Category 9: Smoking Guns (Definitive)

Result: DETECTED / NONE FOUND [Any definitive AI artifacts found]


Weighted Score Calculation

CategoryWeightScoreWeighted
1. Lexical Markers20%XX.XX
2. Statistical Properties15%XX.XX
3. Citation Integrity20%XX.XX
4. Metadiscourse10%XX.XX
5. Structure10%XX.XX
6. Stylometric Features15%XX.XX
7. Voice10%XX.XX
Total100%X.XX/5.0

Confidence Level

[High/Medium/Low -- with modifiers applied]

Assessment

[Score range interpretation -- what collaboration pattern this suggests]

Domain Adjustments Applied

[Any field-specific calibrations]

False Positive Risk Factors

[Any factors that may inflate AI signals: ESL, formal genre, template-driven, etc.]

Actionable Recommendations

[Specific passages or patterns the author should revise to reduce AI-like signals, organized by category. Focus on concrete edits, not vague advice.]


DETECTION FRAMEWORK REFERENCE

The complete framework follows. Apply every element of it.

Comprehensive AI-Generated Text Detection Framework v2.0

Part I: Research Foundation

1.1 Academic Research Sources

This framework synthesizes findings from peer-reviewed research and independent benchmarks:

SourcePublicationKey Contribution
Kobak et al. (2025)*Science Advances*Identified 379 excess vocabulary markers; analyzed 15M+ PubMed abstracts
Liang et al. (2025)*Nature Human Behaviour*Mapped LLM usage across 1M+ papers; 22.5% of CS abstracts show AI modification
Walters & Wilder (2023)*Scientific Reports*Citation hallucination rates (GPT-3.5: 55%, GPT-4: 18%)
RAID Benchmark (2024)*ACL 2024*6M+ generations; comprehensive detector evaluation; FPR-accuracy tradeoffs
Dugan et al. (2024)*COLING 2025 Shared Task*Adversarial robustness testing; detector performance under attack
Stylometric studies*PLOS One*, *Nature H&SS Communications*Function word analysis, POS patterns, phrase structure discrimination

1.2 Commercial Tool Methodologies

ToolPrimary MethodStrengthsLimitations
GPTZeroPerplexity + Burstiness + 7-component MLSentence-level highlighting; educational focusInconsistent on short texts; overflagging reported
TurnitinTransformer deep learning; pattern analysisIntegrated with plagiarism detection; institutional standardFalse positives on formal/ESL writing; institution-only access
CopyleaksML pattern recognition; 100+ languagesMultilingual support; low false positive rates in studiesMixed results on AI-generated content
Originality.aiNeural network trained on AI outputsHigh accuracy on ChatGPT content; fact-checking integrationStruggles with academic essays; pay-per-credit model
PangramActive learning with hard example miningBest adversarial robustness (97.7%); low FPR at strict thresholdsNewer tool; less institutional adoption

1.3 Key Empirical Findings

Detection Accuracy vs. False Positive Rate Tradeoff (RAID 2024)

  • Most detectors achieve high accuracy only at high FPR
  • At FPR <1%, most commercial detectors become ineffective
  • Binoculars method showed best performance at low FPR
  • Adversarial attacks reduce accuracy by 15-40% for most tools

Vocabulary Shift Data (Kobak 2025)

  • "Delve/delving" increased 28x in biomedical literature post-ChatGPT
  • "Underscores" increased 13.8x
  • At least 13.5% of 2024 abstracts were processed with LLMs (lower bound)
  • Effect exceeded even COVID-19 pandemic's vocabulary impact

Stylometric Discrimination (2024-2025 studies)

  • Integrated stylometric features achieve 99%+ discrimination in controlled studies
  • Three most effective features: function word unigrams, POS bigrams, phrase patterns
  • Human raters struggle with AI detection (false positive rates 5%, vs. 1.3% for tools)
  • Humans make judgments based on surface features; stylometry captures deeper patterns

Part II: Evaluation Categories

Category 1: Lexical Markers (Weight: 20%)

Rationale: Kobak et al. (2025) demonstrated that LLMs have distinctive vocabulary preferences that create measurable "excess words" in academic writing.

High-Signal Markers (Frequency Ratio >10x post-ChatGPT)

Word/PhraseFrequency RatioDetection Value
delve/delving28.0xVery High
underscores13.8xVery High
showcasing10.7xVery High
intricateHighHigh
meticulous/meticulouslyHighHigh
multifacetedHighHigh
pivotalHighMedium-High
leveragingHighMedium-High
fosteringHighMedium
nuancedHighMedium
realmHighMedium
groundbreakingHighMedium

Flowery Phrase Patterns

These multi-word constructions strongly indicate AI generation:

  • "meticulously [examining/analyzing/exploring]..."
  • "the intricate [web/tapestry/landscape] of..."
  • "comprehensive [overview/analysis] that delves..."
  • "pivotal role in [fostering/enhancing]..."
  • "navigate the [complex/nuanced] landscape..."
  • "a testament to the [power/importance]..."

Scoring Rubric

ScoreDescription
5Zero high-signal words; natural vocabulary throughout
41-2 medium-signal words in appropriate context
3Multiple medium-signal words OR 1 high-signal word
2Several high-signal words; some flowery phrases
1Pervasive use of AI vocabulary markers; multiple flowery phrases

Category 2: Statistical Properties (Weight: 15%)

Rationale: GPTZero, Turnitin, and academic research consistently identify perplexity and burstiness as core discriminators between AI and human text.

Perplexity (Word Predictability)

LevelCharacteristicIndication
LowHighly predictable word choices; smooth, expected transitionsAI-generated
MediumMixed predictabilityAmbiguous
HighUnexpected word choices; surprising but appropriate vocabularyHuman-written

Human writing indicators:

  • Idiosyncratic vocabulary choices
  • Unexpected but fitting word selections
  • Domain-specific jargon used naturally
  • Personal stylistic preferences evident

Burstiness (Sentence Variation)

LevelCharacteristicIndication
LowUniform sentence lengths; consistent structureAI-generated
MediumSome variation but predictable patternsAmbiguous
HighVariable lengths; mixed structures; rhythm changesHuman-written

Assessment Checklist:

  • Sentence lengths vary substantially (fragments to complex sentences)
  • Paragraph lengths respond to content needs (not uniform)
  • Mix of simple, compound, and complex sentence structures
  • Occasional intentional sentence fragments or run-ons
  • Rhythm changes with content (dense technical to flowing narrative)

Scoring Rubric

ScoreDescription
5High burstiness; highly varied structure; idiosyncratic choices
4Good variation; occasional uniformity in technical sections
3Moderate variation; some predictable patterns
2Low variation; noticeable uniformity
1Very low burstiness; robotic uniformity throughout

Category 3: Citation Integrity (Weight: 20%)

Rationale: Citation hallucination is one of the most definitive markers of AI generation. Walters & Wilder (2023) found 55% fabrication rate in GPT-3.5 and 18% in GPT-4.

Verification Protocol — 100% COVERAGE REQUIRED

Every citation in the paper must be individually verified. For each citation, check:

  1. Author Existence: Do these authors exist and work in this field?
  2. Publication Existence: Does this paper actually exist?
  3. Journal/Venue: Is the journal real? Does it publish this topic?
  4. Date/Volume/Pages: Do the bibliographic details match?
  5. DOI Verification: Does the DOI resolve to the claimed paper?
  6. Claim-Source Alignment: Does the cited source actually support the specific claim made in the paper?

Red Flags for Hallucinated Citations

  • Authors whose names sound plausible but don't exist
  • Papers that combine elements from multiple real sources
  • Journals that don't exist or don't cover the topic
  • Volume/page numbers that don't exist for the claimed year
  • DOIs that lead to different papers or don't resolve
  • Claims that don't match the actual source content

Scoring Rubric

ScoreDescription
5All citations verified as real; all claims accurately represent their sources
4All citations real; minor nuances in claim-source alignment
31-2 problematic citations; most verified
2Multiple fabricated citations or serious misrepresentations
1Pervasive fabrication; many non-existent sources

Category 4: Metadiscourse and Stance Markers (Weight: 10%)

Rationale: Research shows AI text has lower interactional metadiscourse, fewer hedges, and more impersonal tone compared to human academic writing.

Hedging Patterns

Human indicators:

  • Appropriate epistemic caution: "might," "could," "may suggest"
  • Uncertainty acknowledgment: "remains unclear," "further investigation needed"
  • Qualification of claims: "in some cases," "under certain conditions"
  • Personal epistemic markers: "we believe," "it seems to us"

AI indicators:

  • Over-confident assertions
  • Lack of appropriate hedging in speculative claims
  • Generic uncertainty: "more research is needed" without specificity

Boosters and Attitude Markers

Human indicators:

  • Strategic emphasis: "clearly," "importantly," "notably" used judiciously
  • Personal attitude: "surprisingly," "unfortunately," "remarkably"
  • Authorial stance: "we argue," "we contend," "our position is"

AI indicators:

  • Flat emotional expression
  • Absence of authorial stance
  • Generic language without personal investment

Scoring Rubric

ScoreDescription
5Rich metadiscourse; appropriate hedging; clear authorial stance
4Good metadiscourse; some authorial presence
3Adequate hedging but limited stance-taking
2Minimal metadiscourse; generic hedging
1No authorial presence; over-confident or flat tone

Category 5: Structural Characteristics (Weight: 10%)

Rationale: AI text tends toward formulaic organization, uniform paragraph lengths, and predictable structures.

Human Structure Indicators

  • Section organization: Responds to content needs, not template
  • Non-formulaic ordering: May place literature review after intro (field conventions) or integrate throughout
  • Variable paragraph lengths: 1 sentence to 10+ sentences as appropriate
  • Prose-heavy argumentation: Ideas developed in flowing prose, not lists
  • Idiosyncratic organization: Personal approach to presenting material

AI Structure Indicators

  • Formulaic organization: Rigid intro-lit review-methods-results-discussion
  • Uniform paragraph lengths: Consistently 4-6 sentences
  • Heavy list usage: Bullet points and numbered lists as primary format
  • Template adherence: "Tell them what you'll tell them" structure
  • Predictable transitions: "Firstly... Secondly... In conclusion..."

Scoring Rubric

ScoreDescription
5Distinctive organization responding to content; prose-heavy
4Good structural variety; occasional formulaic elements
3Mixed structural signals
2Largely formulaic; heavy list usage
1Rigid template adherence; uniform throughout

Category 6: Stylometric Features (Weight: 15%)

Rationale: Integrated stylometric analysis achieves near-perfect discrimination between human and AI text (Zaitsu & Jin, 2024; Opara, 2024). Three feature categories are most effective: function word unigrams, POS bigrams, and phrase patterns.

6.1 Function Word Patterns

FeatureHuman PatternAI Pattern
Conjunction varietyPersonal preferences (e.g., favors "yet" over "however")Generic, interchangeable usage
Article definitenessConsistent the/a patterns reflecting assumed reader knowledgeInconsistent or overly explicit
Pronoun distributionStable I/we/you ratios appropriate to genreGeneric ratios; "we" as padding
Preposition clusteringNatural collocations (e.g., "in terms of" vs "regarding")Over-reliance on common prepositions
Hedge word preferencesConsistent set (e.g., always "perhaps" not "maybe")Variable, no personal preference

Assessment Method: Sample 5-10 function word choices throughout the document. Do the same choices recur? Does the author seem to have preferences, or are choices interchangeable?

6.2 POS (Part-of-Speech) Patterns

FeatureHuman PatternAI Pattern
Adjective stackingOccasional creative multi-adjective phrasesEither single adjectives or formulaic pairs
Adverb placementVaried (sentence-initial, mid-sentence, end)Predominantly sentence-initial or pre-verb
Verb tense consistencyIntentional shifts for effectRigid consistency or unintentional shifts
Noun phrase complexityVariable (simple to heavily modified)Consistently medium complexity
Subordinate clause densityVaries with content complexityUniform density throughout

Assessment Method: Examine 10 random sentences. Do they show varied syntactic structures, or could they be generated from the same template?

6.3 Phrase Structure Patterns

FeatureHuman PatternAI Pattern
Clause embedding depthVariable (0-3+ levels as needed)Consistently shallow (1-2 levels)
ParallelismIntentional for effect; imperfect elsewhereOver-regular parallel structures
Sentence openingsVaried (subject, adverb, conjunction, subordinate clause)Predominantly subject-first
Rhetorical fragmentsOccasional intentional fragmentsComplete sentences only
List structuresVaried item lengths; occasional incomplete itemsUniform item length and structure

Assessment Method: Examine the first word of 10 consecutive sentences. High variety suggests human authorship; repetitive patterns suggest AI.

6.4 Quantitative Benchmarks

MetricHuman RangeAI Range
Type-Token Ratio (TTR)0.4-0.70.3-0.5
Hapax Legomena Ratio0.4-0.60.2-0.4
Sentence Length CV0.4-0.80.2-0.4

6.5 The "Swap Test"

Could this sentence have been written by any competent author, or does it bear marks of a specific person?

  • If sentences feel interchangeable with generic academic prose -> AI indicator
  • If sentences feel like they could only have been written by this particular author -> Human indicator

Scoring Rubric

ScoreDescription
5Distinctive stylistic fingerprint throughout; passes swap test
4Clear personal style; some generic sections
3Mixed stylistic signals
2Generic style; few distinctive features
1No personal style; template output; fails swap test throughout

Category 7: Voice and Authorial Presence (Weight: 10%)

Indicators of Human Voice

  • Position-taking: Clear arguments, not just summaries
  • Engagement with objections: Anticipates and addresses counterarguments
  • Personal metaphors: Original analogies and comparisons
  • Tonal consistency: Stable voice throughout, appropriate to genre
  • Intellectual curiosity: Questions emerge naturally from argument

Indicators of AI Voice

  • Summary without synthesis: Lists positions without arguing for one
  • Generic comparisons: "Like a blank canvas" type cliches
  • Tonal instability: Shifts between registers
  • Assertion without justification: Claims without reasoning
  • Lack of curiosity: Presents information without wondering

Scoring Rubric

ScoreDescription
5Distinctive intellectual personality throughout; clear voice
4Good authorial presence; consistent tone
3Some voice evident; occasional generic sections
2Weak authorial presence; mostly generic
1No distinctive voice; interchangeable with any author

Category 8: Content Authenticity (optional weight)

Signs of Authentic Scholarship

  • Novel theoretical framework; original synthesis
  • Engagement with contradictions in literature
  • Specific limitations (not generic "more research needed")
  • Research directions that emerge organically from argument
  • Original examples; intellectual risk-taking

Signs of AI Generation

  • Literature as list; summarizes without synthesizing
  • Contradiction avoidance; generic limitations
  • Disconnected future work; familiar examples; safe consensus

Scoring Rubric

ScoreDescription
5Novel synthesis; genuine engagement with tensions; specific contributions
4Good intellectual content; some original insight
3Adequate scholarship; limited novelty
2Primarily summarization; little synthesis
1No original contribution; pure assembly of existing ideas

Category 9: Smoking Guns (Definitive)

If any of these appear, the overall assessment is "AI-generated" regardless of weighted score:

  • "As a large language model..." or similar self-identification
  • "I cannot verify events after my knowledge cutoff..."
  • "Regenerate response" or interface artifacts
  • Embedded instructions or prompt leakage
  • "I don't have access to real-time information..."
  • Responses to hypothetical user queries embedded in text
  • Obvious placeholder text ("[Insert X here]")

Part III: Interpretation Scale

Score RangeAssessment
4.5-5.0Strong evidence of human authorship
4.0-4.4Likely human authorship
3.5-3.9Human with possible AI assistance
3.0-3.4Substantial AI involvement
2.0-2.9Likely AI-generated
1.0-1.9Strong evidence of AI generation

Part IV: Confidence Modifiers

Increase confidence if: Multiple categories converge; citation verification yields concrete evidence; smoking guns detected; long document shows consistent patterns.

Decrease confidence if: Document is short (<500 words); technical/formal writing; non-native English speaker; mixed signals across categories.

Part V: Application Protocol

Step 1: Initial Scan

  1. Check for smoking guns (Category 9)
  2. Run lexical marker check for high-signal words
  3. Assess overall structure and burstiness

If smoking guns detected: Stop -- document is AI-generated. If multiple high-signal markers: Continue with detailed evaluation. If clean initial scan: Continue with full evaluation anyway (thoroughness required).

Step 2: Detailed Evaluation

For each category:

  1. Review indicators and scoring rubric
  2. Document specific evidence from text
  3. Assign score with brief justification

Step 3: Citation Verification (MANDATORY — 100% COVERAGE)

For academic documents, verify every single citation:

  • Confirm the publication exists and authors are real
  • Confirm bibliographic details are accurate
  • Confirm the source supports the specific claim made
  • Report each citation individually in the verification table

Step 4: Synthesis and Assessment

  1. Calculate weighted score
  2. Note convergent/divergent categories
  3. Consider confidence modifiers
  4. Apply domain-specific adjustments
  5. Write summary assessment

Step 5: Actionable Recommendations

Provide specific, concrete revision suggestions organized by category, so the author knows exactly what to change to strengthen the paper against AI-detection signals.

Part VI: Domain-Specific Adjustments

DomainAdjustments
Biomedical/Life SciencesHigher baseline for "comprehensive," "significant findings"; apply Kobak thresholds strictly
Computer ScienceHigher tolerance for technical jargon; watch for code-generation artifacts
HumanitiesExpect higher burstiness; voice category more important
Legal WritingFormal style may mimic AI patterns; focus on citation integrity
Creative WritingVoice and originality most important; structural uniformity less relevant

Part VII: False Positive Risk Factors

  • Non-native English speakers
  • Highly formal writing (legal, regulatory, policy)
  • Template-driven genres (grant proposals, IRB applications)
  • Technical documentation with standardized vocabulary
  • Professionally edited text

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Codex

37.72%
按下载量换算104

Claude

27.67%
按下载量换算76

Cursor

20.39%
按下载量换算56

Gemini CLI

10.22%
按下载量换算28

安全审计

Gen Agent Trust Hub

通过

Socket

通过

Snyk

可疑

权限和风险

需要联网

该 Skill 可能需要联网访问来源站点、仓库或外部 API;具体网络访问范围需要结合源码和 README 复核。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。来源安全扫描存在 warning/failed 结果,不能写成本站确认安全。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills