Token导航 LogoToken导航TokenDH.com
研究检索只读github未标认证来源可访问许可证需确认审计提醒

copilot-history-ingestGitHub Copilot history ingest 搜索

Agent Skill

copilot-history-ingest 用于查找、检索和筛选相关信息,适合在 Codex、Claude、Cursor、Gemini CLI 中需要根据关键词、任务场景或来源线索快速定位候选结果时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

667

周安装

27

GitHub Stars

856

下载量

210
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:copilot-history-ingest(GitHub Copilot history ingest 搜索)
来源仓库:https://github.com/ar9av/obsidian-wiki
仓库路径:skills/copilot-history-ingest
安装命令:
npx skills add https://github.com/ar9av/obsidian-wiki --skill copilot-history-ingest
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

skills.shnpx skills
npx skills add https://github.com/ar9av/obsidian-wiki --skill copilot-history-ingest

简介

用于查找、检索和筛选相关信息。

  • 适合根据关键词、任务场景或来源线索快速定位候选结果。
  • 可结合来源仓库和原始 README 核验具体用法。
  • 安装前建议确认权限范围、维护状态及是否会触发联网或文件读写。
  • copilot-history-ingest 属于研究检索类 Skill,可作为该场景下的辅助能力补充。

SKILL.md

Copilot History Ingest — Conversation Mining

You are extracting knowledge from the user's past GitHub Copilot CLI conversations and distilling it into the Obsidian wiki. Conversations are rich but messy — your job is to find the signal and compile it.

This skill can be invoked directly or via the wiki-history-ingest router (/wiki-history-ingest copilot).

Before You Start

  1. Read .env to get OBSIDIAN_VAULT_PATH and COPILOT_HISTORY_PATH (defaults to ~/.copilot/session-state) and COPILOT_VSCODE_STORAGE_PATH (the VS Code workspaceStorage directory; platform-specific — ask the user if absent from .env)
  2. Read .manifest.json at the vault root to check what's already been ingested
  3. Read index.md at the vault root to know what the wiki already contains

Ingest Modes

Append Mode (default)

Check .manifest.json for each source file (events JSONL, transcript JSONL, checkpoint, session-store DB). Only process:

  • Sessions not in the manifest (new sessions)
  • Sessions whose updated_at is newer than their ingested_at in the manifest

This is usually what you want — the user ran a few new sessions and wants to capture the delta.

Full Mode

Process everything regardless of manifest. Use after a wiki-rebuild or if the user explicitly asks.

GitHub Copilot Data Layout

Copilot stores data in three locations. Scan all three.

Source 1: ~/.copilot/session-state/ (CLI sessions)

~/.copilot/session-state/
├── <session-uuid>/
│   ├── workspace.yaml           # Session metadata (id, cwd, summary_count, created_at, updated_at)
│   ├── vscode.metadata.json     # VS Code context (workspaceFolder, repositoryProperties, customTitle)
│   ├── events.jsonl             # Full event log — all turns, tool calls, reasoning
│   ├── session.db               # Per-session SQLite (todos/todo_deps only — skip for ingestion)
│   ├── index.md                 # Session summary written at session end
│   ├── checkpoints/             # Checkpoint JSON files (mid-session summaries)
│   │   └── <uuid>.json          # title, overview, history, work_done, technical_details,
│   │                            #   important_files, next_steps
│   ├── files/                   # Artifacts produced during session (plans, diagrams, etc.)
│   └── research/                # Research artifacts
└── ...

Source 2: ~/.copilot/session-store.db (Global SQLite)

The canonical cross-session database. This is the highest-value source: structured, queryable, and pre-summarised.

sessions       — id, cwd, repository, branch, summary, created_at, updated_at, host_type
turns          — session_id, turn_index, user_message, assistant_response, timestamp
checkpoints    — session_id, checkpoint_number, title, overview, history, work_done,
                 technical_details, important_files, next_steps, created_at
session_files  — session_id, file_path, tool_name, turn_index, first_seen_at
session_refs   — session_id, ref_type (commit/pr/issue), ref_value, turn_index, created_at
search_index   — FTS5 virtual table (content, session_id, source_type, source_id)

Source 3: VS Code Workspace Storage (<workspaceStorage>/<hash>/GitHub.copilot-chat/)

VS Code extension data, keyed by workspace hash. The path is platform-specific and must come from .env or user input.

<hash>/GitHub.copilot-chat/
├── transcripts/
│   └── <session-uuid>.jsonl     # Conversation transcripts (same JSONL format as events.jsonl)
├── memory-tool/
│   └── memories/
│       └── <base64-session-id>/ # Per-session saved artifacts (plan.md, etc.)
│           └── plan.md
└── codebase-external.sqlite     # Codebase index (skip — no conversation knowledge)

Key data sources ranked by value:

  1. Checkpoints (session-store.db checkpoints table + per-session checkpoints/*.json) — Pre-distilled summaries with overview, work_done, technical_details, important_files, next_steps. Gold.
  2. Session summaries (session-store.db sessions.summary + index.md) — One-paragraph synopsis per session.
  3. Turns (session-store.db turns table + events.jsonl / transcript JSONL) — Full conversation. Rich but verbose.
  4. Memory artifacts (memory-tool/memories/<id>/plan.md etc.) — Pre-written plans and structured notes the user saved explicitly. Worth importing verbatim (or lightly summarised).
  5. File access patterns (session_files table + tool.execution_* events) — Which files the agent repeatedly touched — reveals high-value project files.
  6. Session refs (session_refs table) — Commits, PRs, and issues linked to sessions.
  7. vscode.metadata.json — Workspace folder path, branch, customTitle (user-set session label). Useful for grouping and naming.

Step 1: Survey and Compute Delta

Scan all three data locations and compare against .manifest.json:

# --- Source 1: per-session directories ---
# Find all session directories (each has workspace.yaml)
ls ~/.copilot/session-state/

# For each session, read workspace.yaml for id/cwd/updated_at
# and vscode.metadata.json for customTitle / repositoryProperties

# --- Source 2: global database ---
# Query session-store.db with sqlite3 (or Python sqlite3)
SELECT s.id, s.cwd, s.repository, s.branch, s.summary, s.updated_at,
       COUNT(DISTINCT t.turn_index) AS turn_count,
       COUNT(DISTINCT c.id)         AS checkpoint_count
FROM sessions s
LEFT JOIN turns t ON t.session_id = s.id
LEFT JOIN checkpoints c ON c.session_id = s.id
GROUP BY s.id
ORDER BY s.updated_at DESC;

# --- Source 3: VS Code workspace storage ---
# For each <hash> directory under workspaceStorage, check for GitHub.copilot-chat/
# Find transcript files
ls <workspaceStorage>/<hash>/GitHub.copilot-chat/transcripts/

Build a unified inventory — one entry per session UUID — and classify:

  • New — not in manifest → needs ingesting
  • Modified — in manifest but updated_at is newer → needs re-ingesting
  • Unchanged — in manifest and not modified → skip in append mode

Report to the user: "Found X sessions in session-state, Y in session-store.db, Z VS Code transcript files. Checkpoints: A. Delta: B new, C modified."

Step 2: Ingest Checkpoints and Summaries First

Checkpoints are already distilled — process them before touching raw turns.

From session-store.db:

SELECT s.id, s.cwd, s.repository, s.branch, s.summary,
       c.checkpoint_number, c.title, c.overview, c.work_done,
       c.technical_details, c.important_files, c.next_steps,
       c.created_at
FROM checkpoints c
JOIN sessions s ON c.session_id = s.id
ORDER BY s.updated_at DESC, c.checkpoint_number ASC;

From per-session checkpoints/*.json:

Each checkpoint file has: title, overview, history, work_done, technical_details, important_files, next_steps.

Read index.md (if present) as a session-level summary — it's typically written at session end and is already concise.

What to extract:

  • overview → high-level description of what the session accomplished
  • work_done → concrete tasks completed (good for skills / project pages)
  • technical_details → implementation specifics (good for concepts pages)
  • important_files → high-value files in the project (good for project pages)
  • next_steps → open threads (good for linking to ongoing project work)

Step 3: Parse Session Turns

Read turns from session-store.db (preferred — already parsed) or from events.jsonl / transcript JSONL.

From session-store.db:

SELECT turn_index, user_message, assistant_response, timestamp
FROM turns
WHERE session_id = '<uuid>'
ORDER BY turn_index ASC;

From events.jsonl / transcript JSONL:

Each file is one session. Each line is a JSON event. See references/copilot-data-format.md for the full schema.

Relevant event types:

typeWhat it isWorth reading?
session.startSession metadata (cwd, branch, version)Yes — establishes project context
user.messageUser turnYes — data.content
assistant.messageAssistant turnYes — data.content (text) + data.toolRequests
tool.execution_startTool callSkim — reveals what files/commands were used
tool.execution_endTool resultNo — usually noise

Extraction strategy for assistant.message:

  • data.content is the assistant's text response — extract this
  • data.reasoningText is internal reasoning — skip (it's the unpacked reasoningOpaque field)
  • data.toolRequests lists tool calls — skim tool names and arguments for file access patterns
  • Skip type: "tool.execution_end" entirely

Step 3b: Process Memory Artifacts

For each session that has a memory-tool/memories/<base64-id>/ directory in VS Code workspace storage, read any markdown files saved there (typically plan.md). These are documents the user explicitly saved — treat them as high-quality, user-authored content.

Decode the base64 directory name to get the session UUID:

import base64
session_id = base64.b64decode(dir_name).decode('utf-8')

Memory artifacts map to project skills/ or concepts/ pages, depending on content type.

Step 3c: Extract File and Ref Patterns

From session-store.db:

-- Most-touched files per project
SELECT repository, file_path, COUNT(*) AS touch_count
FROM session_files
GROUP BY repository, file_path
ORDER BY touch_count DESC;

-- Linked commits/PRs/issues per session
SELECT session_id, ref_type, ref_value, turn_index
FROM session_refs
ORDER BY session_id, turn_index;

File access patterns reveal which files are architecturally important — note them on project pages.

Session refs link Copilot sessions to git history — useful for connecting wiki knowledge to concrete code changes.

Step 4: Cluster by Topic

Don't create one wiki page per session. Instead:

  • Group extracted knowledge by topic across sessions
  • A single session about "debugging auth + setting up CI" → two separate topics
  • Three sessions across different days about "React performance" → one merged topic
  • cwd / repository give you a natural first-level grouping; vscode.metadata.json's customTitle gives a human-readable session label

Step 5: Distill into Wiki Pages

Each Copilot project maps to a project directory in the vault. Derive the project name from cwd or repository:

C:\Users\name\git\my-project   → my-project
/Users/name/code/another-app   → another-app

Prefer repository (e.g., owner/repo) from session-store.db over raw cwd when available.

Project-specific vs. global knowledge

What you foundWhere it goesExample
Project architecture decisionsprojects/<name>/concepts/projects/my-project/concepts/main-architecture.md
Project-specific debugging patternsprojects/<name>/skills/projects/my-project/skills/api-rate-limiting.md
General concept the user learnedconcepts/ (global)concepts/react-server-components.md
Recurring problem across projectsskills/ (global)skills/debugging-hydration-errors.md
A tool/service usedentities/ (global)entities/vercel-functions.md
Patterns across many sessionssynthesis/ (global)synthesis/common-debugging-patterns.md

For each project with content, create or update the project overview page at projects/<name>/<name>.mdnamed after the project, not _project.md. Obsidian's graph view uses the filename as the node label, so _project.md makes every project show up as _project in the graph. Naming it <name>.md gives each project a distinct, readable node name.

Important: Distill the *knowledge*, not the conversation. Don't write "In a session on March 15, the user asked about X." Write the knowledge itself, with the session as a source attribution.

Write a summary: frontmatter field on every new/updated page — 1–2 sentences, ≤200 chars, answering "what is this page about?" for a reader who hasn't opened it. wiki-query's cheap retrieval path reads this field to avoid opening page bodies.

Mark provenance per the convention in llm-wiki (Provenance Markers section):

  • Checkpoints and index.md are pre-distilled by the system — treat checkpoint-derived claims as extracted (the system wrote them from observed actions).
  • Memory artifacts are user-authored — treat as extracted.
  • Conversation turn distillation is mostly inferred. You're synthesizing a coherent claim from many turns. Apply ^[inferred] liberally to synthesized patterns, generalizations across sessions, and "what the user really meant" interpretations.
  • Use ^[ambiguous] when the user changed direction mid-session or when the session ended unresolved.
  • Write a provenance: frontmatter block on every new/updated page summarizing the rough mix.

Step 6: Update Manifest, Journal, and Special Files

Update .manifest.json

For each session processed, add/update its entry with:

  • ingested_at, session_id, updated_at
  • source_type: one of "copilot_session", "copilot_checkpoint", "copilot_transcript", "copilot_memory_artifact"
  • project: the decoded project name
  • pages_created and pages_updated lists

Also update the projects section of the manifest:

{
  "project-name": {
    "repository": "owner/repo",
    "cwd": "C:\\Users\\name\\git\\project-name",
    "vault_path": "projects/project-name",
    "last_ingested": "TIMESTAMP",
    "sessions_ingested": 5,
    "sessions_total": 8,
    "checkpoints_ingested": 12,
    "memory_artifacts_ingested": 3
  }
}

Create journal entry + update special files

Update index.md and log.md per the standard process:

- [TIMESTAMP] COPILOT_HISTORY_INGEST projects=N sessions=M checkpoints=C pages_updated=X pages_created=Y mode=append|full

hot.md — Read $OBSIDIAN_VAULT_PATH/hot.md (create from the template in wiki-ingest if missing). Update Recent Activity with a one-line summary — e.g. "Ingested 5 Copilot sessions across 2 projects; surfaced patterns in API design and testing strategy." Keep the last 3 operations. Update Active Threads if any ongoing project is now better understood. Update updated timestamp.

Privacy

  • Distill and synthesize — don't copy raw conversation text verbatim
  • Skip anything that looks like secrets, API keys, passwords, tokens
  • data.reasoningOpaque / data.reasoningText in assistant events is internal reasoning — skip entirely, never copy to wiki
  • If you encounter personal/sensitive content, ask the user before including it
  • The user's conversations may reference other people — be thoughtful about what goes in the wiki

Reference

See references/copilot-data-format.md for detailed data structure documentation.

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Codex

36.79%
按下载量换算77

Claude

30.83%
按下载量换算65

Cursor

20.25%
按下载量换算43

Gemini CLI

8.74%
按下载量换算18

安全审计

Gen Agent Trust Hub

可疑

Socket

通过

Snyk

通过

权限和风险

只读

该 Skill 主要提供规则、说明或参考内容,本身偏只读;真正读写文件、联网或执行命令仍取决于宿主 Agent 的任务。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。来源安全扫描存在 warning/failed 结果,不能写成本站确认安全。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills