swain-search
Collect, normalize, and cache source materials into reusable troves that swain-design artifacts can reference.
Mode detection
| Signal | Mode |
|---|---|
| No trove exists for the topic, or user says "research X" / "gather sources" | Create — new trove |
| Trove exists and user provides new sources or says "add to" / "extend" | Extend — add sources to existing trove |
| Trove exists and user says "refresh" or sources are past TTL | Refresh — re-fetch stale sources |
| User asks "what troves do we have" or "find sources about X" | Discover — search existing troves by tag |
Prior art check
Before creating a new trove or running web searches, scan existing troves for relevant content. This avoids duplicating research and surfaces connections to prior work.
Phase 1 — Literal keyword match
Search for the source name, URL fragments, and author name:
# Search trove manifests by tag
grep -rl "<keyword>" docs/troves/*/manifest.yaml 2>/dev/null
# Search trove source content
grep -rl "<keyword>" docs/troves/*/sources/**/*.md 2>/dev/null
# Search trove syntheses
grep -rl "<keyword>" docs/troves/*/synthesis.md 2>/dev/nullPhase 2 — Semantic topic match
After fetching the source and understanding what it's about, extract 3-5 topic keywords from the source's *content* (not just its name or URL). Then search existing troves by topic:
# Search trove tags for topic keywords
grep -l "<topic-keyword-1>\|<topic-keyword-2>\|<topic-keyword-3>" docs/troves/*/manifest.yaml 2>/dev/null
# Search synthesis summaries for topic keywords
grep -l "<topic-keyword-1>\|<topic-keyword-2>\|<topic-keyword-3>" docs/troves/*/synthesis.md 2>/dev/nullTopic keywords should describe what the source *is about*, not what it's *called*. For example, a repo named "Cog" that implements a memory system for Claude Code should generate topic keywords like agent-memory, memory-architecture, claude-code, persistent-memory — not cog or marciopuga.
If the source has not been fetched yet (URL-only invocation), use whatever topic information is available from the URL or title and defer full topic matching until after the source is fetched.
Decision gate
Before proceeding to Create or Extend mode, output a visible routing decision:
Prior art check: Phase 1 found [N matches / no matches]. Phase 2 found [N matches / no matches]: [trove-id (tags: x, y),...]. Decision: Extending [trove-id] / Creating new trove [slug] because [reason].
This makes the trove routing decision auditable. If any trove matches on 2+ topic keywords, default to Extend mode unless the topic is genuinely distinct (adjacent but different subject matter).
Action on matches
If existing troves contain relevant sources:
- Report what was found — show the trove ID, matching source titles, and relevant excerpts
- Suggest extend over create — if an existing trove covers the same topic, extend it rather than creating a parallel trove
- Cross-link — if the topic is adjacent but distinct, create a new trove but note the related trove in synthesis.md
This step runs in all modes (Create, Extend, Discover) and before any web searches. Existing trove content is always checked first.
Snapshot evidence gate (SPEC-220)
Before a remote source can be treated as collected evidence, the run must produce a raw snapshot and a metadata ledger entry in .agents/search-snapshots/metadata.jsonl.
Required flow for remote sources:
- Export/download the raw snapshot first:
- bash skills/swain-search/scripts/export-snapshot.sh --url "<source-url>" --out-dir ".agents/search-snapshots/raw"
- Normalize the downloaded file using
writing-skillsorskill-creator(never summary-only browser notes). - Log metadata:
- bash skills/swain-search/scripts/log-snapshot-metadata.sh --source-url "<source-url>" --export-mode "<mode>" --raw-path "<raw-path>" --normalized-path "<normalized-path>" --normalization-skill "<writing-skills|skill-creator>"
- Verify before publication:
- bash skills/swain-search/scripts/verify-snapshot-evidence.sh --source-url "<source-url>"
If verification fails, mark the source unverified, do not publish it downstream, and report the warning to the operator.
Create mode
Build a new trove from scratch.
Step 1 — Gather inputs
Ask the user (or infer from context) for:
- Trove ID — a slug for the topic (e.g.,
websocket-vs-sse). Suggest one if the context is clear. - Tags — keywords for discovery (e.g.,
real-time,websocket,sse) - Sources — any combination of:
- Web search queries ("search for WebSocket vs SSE comparisons") - URLs (web pages, forum threads, docs) - Video/audio URLs - Local file paths
- Freshness TTL overrides — optional, defaults are fine for most troves
If invoked from swain-design (e.g., spike entering Active), the artifact context provides the topic, tags, and sometimes initial sources.
Step 2 — Collect and normalize
For each source, use the appropriate capability. Read references/normalization-formats.md for the exact markdown structure per source type.
Web search queries:
- Use a web search capability to find relevant results
- Select the top 3-5 most relevant results
- For each: fetch the page, normalize to markdown per the web page format
- If no web search capability is available, tell the user and skip
Web page URLs:
- Fetch the page using a browser or page-fetching capability
- Strip boilerplate (nav, ads, sidebars, cookie banners)
- Normalize to markdown per the web page format
- If fetch fails, record the URL in manifest with a
failed: trueflag and move on
Google Docs / Drive-like documents:
- Export raw content first (required):
- bash skills/swain-search/scripts/export-snapshot.sh --url "<source-url>" --out-dir ".agents/search-snapshots/raw"
- Prefer API export modes (
google-doc-export,google-slides-export,google-drive-download). - If API export fails, use a browser helper fallback only when available.
- Normalize the exported file with
writing-skillsorskill-creator. - Log metadata in
.agents/search-snapshots/metadata.jsonl. - Verify with
verify-snapshot-evidence.shbefore including the source in trove outputs.
Paywall proxy fallback:
After fetching a web page, check if a paywall proxy is available for the URL's domain:
- Run
scripts/resolve-proxy.sh <url>
- Exit 1: no proxy configured — use the direct fetch content as-is - Exit 0: outputs PROXY:<name>:<proxy-url> and SIGNAL:<text> lines
- If exit 0, check the fetched content for each
SIGNALtext (case-sensitive literal match) - If any signal matches (or the article body is under ~200 words):
- Log: "Paywall detected for <url> — trying proxy fallback" - Try each PROXY URL in order, fetching via the same page-fetching capability used for web pages - First proxy that returns substantive content (more than the truncated original) wins - Set proxy-used: <name> and notes: "Full article retrieved via <name> proxy" in the manifest entry
- If no signals match: use the direct fetch content as-is (no proxy needed)
- If all proxies fail: keep the original truncated content, set
notes: "Paywalled; proxies exhausted — content from direct fetch only"
The registry lives at references/paywall-proxies.yaml. Add new domains or proxies there — no skill file changes needed.
Video/audio URLs:
- Use a media transcription capability to get the transcript
- Normalize to markdown per the media format (timestamps, speaker labels, key points)
- If no transcription capability is available, tell the user and skip — or accept a pre-made transcript
Local files:
- Use a document conversion capability (PDF, DOCX, etc.) or read directly if already markdown
- Normalize per the document format using
writing-skillsorskill-creator - For markdown files: add frontmatter only, preserve content
Forum threads / discussions:
- Fetch and normalize per the forum format (chronological, author-attributed)
- Flatten nested threads to chronological order with reply-to context
Repositories:
- Clone or read the repository contents
- Mirror the original directory tree under
sources/<source-id>/ - Default: mirror the full tree. For large repositories (thousands of files), ingest selectively and set
selective: truein the manifest entry - Populate the
highlightsarray with paths to the most important files (relative to the source-id directory)
Documentation sites:
- Crawl or fetch the documentation site
- Mirror the section hierarchy under
sources/<source-id>/ - Default: mirror the full site. For large sites, ingest selectively and set
selective: true - Populate the
highlightsarray with paths to the most important pages - Preserve internal link structure where possible
Each normalized source gets a slug-based source ID and lives in a directory-per-source layout:
- Flat sources (web, forum, media, document, local):
sources/<source-id>/<source-id>.md - Hierarchical sources (repository, documentation-site):
sources/<source-id>/with the original tree mirrored inside
Source ID generation:
- Derive the source ID as a slug from the source title or URL (e.g.,
mdn-websocket-api,strangeloop-2025-realtime) - When a slug collides with an existing source ID: append
__word1-word2using two random words fromreferences/wordlist.txt - If the wordlist is missing, append
__followed by 4 hex characters (e.g.,__a3f8) as a fallback
Step 3 — Generate manifest
Create manifest.yaml following the schema in references/manifest-schema.md. Include:
- Trove metadata (id, created date, tags)
- Default freshness TTL per source type
- One entry per source with provenance (URL/path, fetch date, content hash, type)
Compute content hashes as bare hex SHA-256 digests (no prefix) of the normalized markdown content:
shasum -a 256 sources/mdn-websocket-api/mdn-websocket-api.md | cut -d' ' -f1Step 4 — Generate synthesis
Create synthesis.md — a structured distillation of key findings across all sources.
Structure the synthesis by theme, not by source. Group related findings together, cite sources by ID, and surface:
- Key findings — what the sources collectively say about the topic
- Points of agreement — where sources converge
- Points of disagreement — where sources conflict or present alternatives
- Gaps — what the sources don't cover that might matter
Keep it concise. The synthesis is a starting point, not a comprehensive report — the user or artifact author will refine it.
Step 5 — Commit and stamp
Use the dual-commit pattern (same as swain-design lifecycle stamps) to give the trove a reachable commit hash.
Before Commit A — append a history entry to manifest.yaml with a -- placeholder for the commit hash:
history:
- event: created
date: 2026-03-09
commit: "--"
sources: 3Commit A — commit the trove content:
git add docs/troves/<trove-id>/
git commit -m "research(<trove-id>): create trove with N sources"
TROVE_HASH=$(git rev-parse HEAD)Commit B — back-fill the commit hash into the history entry, then update the referencing artifact's frontmatter (if one exists):
# Replace "--" with the real hash in the history entry
# Update artifact frontmatter: trove: <trove-id>@<TROVE_HASH>
git add docs/troves/<trove-id>/manifest.yaml
git add docs/<artifact-type>/<phase>/<artifact-dir>/ # if artifact exists
git commit -m "docs(<trove-id>): stamp history hash ${TROVE_HASH:0:7}"If no referencing artifact exists yet (standalone research), Commit B still stamps the history entry — report the hash so it can be referenced later.
Push — after Commit B, push to origin/trunk so the trove is immediately available to other agents and sessions:
git push origin trunkStep 6 — Report
Tell the user what was created:
Trove<trove-id>created with N sources — committed as<TROVE_HASH:0:7>. -docs/troves/<trove-id>/manifest.yaml— provenance and metadata -docs/troves/<trove-id>/sources/— N normalized source files -docs/troves/<trove-id>/synthesis.md— thematic distillation Reference from artifacts with:trove: <trove-id>@<TROVE_HASH:0:7>
Extend mode
Add new sources to an existing trove.
- Read the existing
manifest.yaml - Collect and normalize new sources (same as Create step 2)
- Assign slug-based source IDs to new sources (following the same ID generation rules)
- Append new entries to
manifest.yaml - Update
refresheddate - Regenerate
synthesis.mdincorporating all sources (old + new) - Append a
historyentry withevent: extendedandcommit: "--"placeholder - Commit and stamp (same dual-commit pattern as Create step 5):
- Commit A: git commit -m "research(<trove-id>): extend with N new sources" - Capture TROVE_HASH=$(git rev-parse HEAD) - Commit B: back-fill hash in history entry, update referencing artifact frontmatter (if artifact exists) - Push: git push origin trunk
- Report what was added, including the new commit hash
Refresh mode
Re-fetch stale sources and update changed content.
- Read
manifest.yaml - For each source, check if
fetcheddate +freshness-ttlhas elapsed - For stale sources:
- Re-fetch the raw content - Re-normalize to markdown - Compute new content hash - If hash changed: replace the source file, update manifest entry - If hash unchanged: update only fetched date
- Update
refresheddate in manifest - If any content changed, regenerate
synthesis.md - Append a
historyentry withevent: refreshed,sources-changed: M, andcommit: "--"placeholder - Commit and stamp (same dual-commit pattern as Create step 5):
- Commit A: git commit -m "research(<trove-id>): refresh N sources (M changed)" - Capture TROVE_HASH=$(git rev-parse HEAD) - Commit B: back-fill hash in history entry, update referencing artifact(s) frontmatter — check referenced-by in manifest for all dependents - Push: git push origin trunk
- Report: "Refreshed N sources. M had changed content, K were unchanged. New hash:
<TROVE_HASH:0:7>."
For sources with freshness-ttl: never, skip them during refresh.
Discover mode
Help the user find existing troves relevant to their topic.
- Scan
docs/troves/*/manifest.yamlfor all troves - Match against the user's query by:
- Tag match — trove tags contain query keywords - Title match — trove ID slug contains query keywords
- For each match, show: trove ID, tags, source count, last refreshed date, referenced-by list
- If no matches, suggest creating a new trove
Graceful degradation
The skill references capabilities generically. When a capability isn't available:
| Capability | Fallback |
|---|---|
| Web search | Skip search-based sources. Tell user: "No web search capability available — provide URLs directly or add a search MCP." |
| Browser / page fetcher | Try basic URL fetch. If that fails: "Can't fetch this URL — paste the content or provide a local file." |
| Snapshot export for remote docs | If export fails and no helper exists: mark source unverified, do not publish downstream, report exact URL and failure mode. |
| Media transcription | "No transcription capability available — provide a pre-made transcript file, or add a media conversion tool." |
| Document conversion | "Can't convert this file type — provide a markdown version, or add a document conversion tool." |
| Paywall proxy | Keep truncated content. Note in manifest: "Paywalled; proxies exhausted." Suggest user provide content manually. |
Never fail the entire run because one capability is missing. Collect what you can, skip what you can't, and report clearly.
Capability detection
Before collecting sources, check what's available. Look for tools matching these patterns — the exact tool names vary by installation:
- Web search: tools with "search" in the name (e.g.,
brave_web_search,bing-search-to-markdown) - Page fetching: tools with "fetch", "webpage", "browser" in the name (e.g.,
fetch_content,webpage-to-markdown,browser_navigate) - Media transcription: tools with "audio", "video", "youtube" in the name (e.g.,
audio-to-markdown,youtube-to-markdown) - Document conversion: tools with "pdf", "docx", "pptx", "xlsx" in the name (e.g.,
pdf-to-markdown,docx-to-markdown)
Report available capabilities at the start of collection so the user knows what will and won't work.
Linking from artifacts
Artifacts reference troves in frontmatter:
trove: websocket-vs-sse@abc1234The format is <trove-id>@<commit-hash>. The commit hash pins the trove to a specific version — troves evolve over time as sources are added or refreshed, and the hash ensures reproducibility.
The dual-commit workflow in Create step 5, Extend step 8, and Refresh step 7 handles this automatically — Commit A records the trove content and Commit B stamps the hash into the history entry and referencing artifact's frontmatter. Do not defer this to the operator.