Token导航 LogoToken导航TokenDH.com
研究检索需要联网github未标认证来源可访问许可证需确认审计提醒

dev-ai-coding-metrics开发 AI 编码指标

Agent Skill

dev-ai-coding-metrics 用于查找、检索和筛选相关信息,适合在 Codex、Claude、Cursor、Gemini CLI 中需要根据关键词、任务场景或来源线索快速定位候选结果时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

984

周安装

41

GitHub Stars

59

下载量

328
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:dev-ai-coding-metrics(开发 AI 编码指标)
来源仓库:https://github.com/vasilyu1983/ai-agents-public
仓库路径:skills/dev-ai-coding-metrics
安装命令:
npx skills add https://github.com/vasilyu1983/ai-agents-public --skill dev-ai-coding-metrics
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

skills.shnpx skills
npx skills add https://github.com/vasilyu1983/ai-agents-public --skill dev-ai-coding-metrics

简介

用于查找、检索和筛选相关信息,适合快速定位候选结果。

  • 可结合关键词、任务场景或来源线索进行信息聚合。
  • 通过命令行安装并使用,需参考原始 README 获取具体指令。
  • 安装前建议确认是否会触发联网或文件读写,确保权限可控。
  • dev-ai-coding-metrics 属于研究检索类 Skill,可作为该场景下的辅助能力补充。

SKILL.md

AI Coding Agent Metrics for Engineering Teams

Measure what matters when adopting AI coding tools. This skill provides metrics frameworks, measurement methodology, ROI models, and reporting templates for engineering managers, VPs of Engineering, and CTOs evaluating or scaling AI coding agents.

When to Use This Skill

  • Evaluating AI coding tool ROI before or after purchase
  • Building a metrics program for AI-assisted development
  • Reporting AI tool impact to leadership or board
  • Designing controlled experiments to measure AI effectiveness
  • Comparing productivity across AI-equipped and traditional teams
  • Tracking adoption health and identifying stall patterns
  • Assessing quality impact of AI-generated code
  • Running developer experience surveys for AI tools

Quick Reference

TaskReferenceAsset
Track tool adoptionadoption-metrics.mdadoption-survey-template.md
Measure productivityproductivity-metrics.mdmetric-dashboard-template.md
Monitor code qualityquality-metrics.mdmetric-dashboard-template.md
Calculate ROIroi-framework.mdroi-calculator-template.md
Assess developer experiencedeveloper-experience-metrics.mdadoption-survey-template.md
Design experimentsbenchmarking-methodology.mdexperiment-design-template.md
Report to executivesroi-framework.mdexecutive-report-template.md
Measure AI coding impactthis skill
Context engineering for AIdev-context-engineering
Per-task agent ROIai-agents
Observability for systemsqa-observability

Core Metrics Taxonomy

Five measurement categories. Start with Adoption (you can't optimize what people aren't using), then layer in the others.

1. Adoption Metrics

Track whether and how developers use AI tools.

MetricFormulaTarget (Mature)Source
License Utilizationactive_users / licensed_seats>85%License admin
DAU/WAU Ratiodaily_active / weekly_active>0.6Tool telemetry
Feature Breadthfeatures_used / features_available>0.5Tool telemetry
Acceptance Ratesuggestions_accepted / suggestions_shown25-35%Copilot API / tool logs
Organic Usage Ratiovoluntary_sessions / total_sessions>0.8Survey + telemetry

Deep dive: references/adoption-metrics.md — 8 additional metrics, adoption curve phases, tool-specific tracking, stall patterns.

2. Velocity Metrics

Measure speed and throughput changes.

MetricFormulaExpected AI ImpactSource
Deploy Frequencydeploys / time_period+15-30%CI/CD pipeline
Lead Time for Changescommit_to_production-20-40%Git + CI/CD
Cycle Timeticket_start_to_deploy-15-35%Project management + Git
PR Throughputmerged_PRs / developer / week+20-40%Git platform
Time to First Commitonboard_date_to_first_commit-30-50%Git + HR data

Deep dive: references/productivity-metrics.md — DORA adaptations, SPACE framework, cycle time decomposition, confounding variables.

3. Quality Metrics

Track whether AI helps or hurts code quality.

MetricFormulaWatch DirectionSource
Bug Densitybugs / KLOCShould decreaseIssue tracker
Defect Escape Rateprod_bugs / total_bugsShould decreaseIssue tracker
Rework Ratefollowup_PRs / total_PRsWatch for increaseGit platform
Test Coveragecovered_lines / total_linesShould increaseCI coverage
Vulnerability Ratenew_vulns / sprintWatch for increaseSAST tools

Critical warning: Early studies show mixed quality results. AI can increase velocity while *also* increasing bug density if guardrails are missing. Monitor both.

Deep dive: references/quality-metrics.md — complexity tracking, security metrics, technical debt, quality guardrails.

4. Economic Metrics

Calculate costs, benefits, and ROI.

MetricFormulaBenchmarkSource
Cost per Seat(license + infra + training) / developers$20-50/dev/monthFinance
Hours Saved/Dev/Weekmeasured_or_estimated_time_savings2-8 hrs (varies widely)Survey + telemetry
ROI(net_benefits - costs) / costs × 100100-300% yr1 (vendor data)Calculated
Payback Periodtotal_investment / monthly_net_benefit2-6 monthsCalculated
Break-Even Adoptioncost / (max_benefit × developers)25-40% of teamCalculated

Caveat: Most published ROI figures come from tool vendors. Independent studies show lower but still positive returns. Always triangulate.

Deep dive: references/roi-framework.md — cost model, value model, formulas, executive reporting, benchmarks with caveats.

5. Experience Metrics

Measure developer satisfaction and cognitive impact.

MetricFormulaTargetSource
AI Tool Satisfactionsurvey_score (1-5 Likert)>3.8/5.0Quarterly survey
Tool NPSpromoters% - detractors%>30Quarterly survey
Cognitive LoadNASA-TLX adaptation (1-7)<4.0/7.0Post-task survey
Give-Up Ratestarted_AI_finished_manual / total<20%Telemetry
Trust Calibrationappropriate_review_rate>80%Code review data

Deep dive: references/developer-experience-metrics.md — survey design, cognitive load measurement, friction indicators, trust metrics.


Measurement Maturity Model

Where is your organization in measuring AI coding impact?

LevelNameCharacteristicsKey Action
L0No MeasurementNo tracking beyond license countInstall basic telemetry, run first survey
L1Basic TrackingLicense utilization + adoption rate trackedAdd DORA metrics baseline, first ROI estimate
L2Structured ProgramDORA + adoption + quality metrics active, quarterly surveyDesign controlled experiment, build dashboard
L3Evidence-BasedControlled experiments, statistical rigor, executive reportingCross-team benchmarking, predictive models
L4OptimizedContinuous measurement, automated dashboards, data-driven tool selectionIndustry benchmarking, publish findings

L0 → L1 Quick Start (2 hours)

  1. Pull license utilization from admin console
  2. Run the adoption survey (assets/adoption-survey-template.md)
  3. Calculate basic ROI estimate (assets/roi-calculator-template.md)
  4. Present 1-page summary to leadership (assets/executive-report-template.md)

L1 → L2 (2-4 weeks)

  1. Establish DORA metric baselines (references/productivity-metrics.md)
  2. Set up quality tracking (references/quality-metrics.md)
  3. Build three-tier dashboard (assets/metric-dashboard-template.md)
  4. Schedule quarterly developer experience surveys

L2 → L3 (1-3 months)

  1. Design first controlled experiment (assets/experiment-design-template.md)
  2. Apply statistical rigor (references/benchmarking-methodology.md)
  3. Create executive reporting cadence (assets/executive-report-template.md)
  4. Cross-reference with dev-context-engineering maturity model for context quality impact

L3 → L4 (ongoing)

  1. Automate data collection and dashboards
  2. Build predictive models (adoption → productivity correlation)
  3. Benchmark against industry data
  4. Contribute findings to community (conference talks, blog posts)

Metric Selection Decision Tree

Not every org needs every metric. Start from what you're trying to prove.

WHAT ARE YOU TRYING TO PROVE?
  │
  ├─ "Should we buy AI coding tools?"
  │   └─ START: roi-framework.md → roi-calculator-template.md
  │       Metrics: cost per seat, estimated hours saved, break-even adoption rate
  │
  ├─ "Are developers actually using the tools?"
  │   └─ START: adoption-metrics.md → adoption-survey-template.md
  │       Metrics: DAU/WAU, acceptance rate, feature breadth, organic usage
  │
  ├─ "Are we shipping faster?"
  │   └─ START: productivity-metrics.md → metric-dashboard-template.md
  │       Metrics: DORA metrics, cycle time, PR throughput
  │
  ├─ "Is code quality suffering?"
  │   └─ START: quality-metrics.md → metric-dashboard-template.md
  │       Metrics: bug density, defect escape rate, rework rate, vulnerability rate
  │
  ├─ "Are developers happy with AI tools?"
  │   └─ START: developer-experience-metrics.md → adoption-survey-template.md
  │       Metrics: satisfaction, NPS, cognitive load, give-up rate
  │
  ├─ "How do we compare to industry?"
  │   └─ START: benchmarking-methodology.md → experiment-design-template.md
  │       Metrics: DORA benchmarks, adoption curves, ROI ranges
  │
  └─ "Should we expand or cut the program?"
      └─ COMBINE: roi-framework.md + adoption-metrics.md + executive-report-template.md
          Metrics: ROI trend, adoption trajectory, satisfaction trend, quality delta

Dashboard Design Principles

Three-Tier Hierarchy

TierAudienceRefreshMetricsPurpose
ExecutiveC-Suite, VP EngMonthly4-6 KPIsInvestment decision, program health
Team LeadEng ManagersWeekly8-10 metricsTeam optimization, coaching
DeveloperIndividual devsReal-timePersonal statsSelf-improvement (opt-in only)

Design Rules

  1. Lead with outcomes, not activity — show deploy frequency, not lines of code
  2. Always show trend lines — a single number is meaningless without direction
  3. Include confidence indicators — mark metrics with low sample sizes or high variance
  4. Never rank individuals — aggregate to team level minimum (team size ≥5)
  5. Pair speed with quality — never show velocity without adjacent quality metrics
  6. Show cost alongside benefit — ROI is a ratio, not a cherry-picked benefit number

See: assets/metric-dashboard-template.md for full layout.


Anti-Patterns

Anti-PatternWhy It's HarmfulFix
Lines of Code as productivityAI inflates LOC; rewards verbosity over clarityUse outcome metrics (features shipped, bugs resolved)
Individual developer trackingCreates surveillance culture, erodes trustAggregate to team level, minimum team size 5
Vanity metrics only"90% adoption!" means nothing if output quality dropsAlways pair adoption with quality and satisfaction
Measuring too earlyFirst 4 weeks are learning curve, not steady stateAllow 8-12 week adoption curve before measuring impact
Vendor benchmarks as gospelVendor studies select favorable conditionsTriangulate with independent research; discount vendor data 30-50%
Ignoring the denominator"Shipped 40% more PRs" — but were they smaller?Normalize metrics (features/sprint, not PRs/sprint)
Correlation → causationTeam adopted AI *and* got a new senior devUse controlled experiments (benchmarking-methodology.md)
Surveying without actingDevelopers report friction → nothing changesClose the loop: share results + action plan within 2 weeks
One metric to rule them allSingle metric always gets gamedUse balanced scorecard (adoption + velocity + quality + experience)
Comparing incomparable teamsFrontend team vs infra team → meaningless comparisonSegment by project type, stack, and task complexity

Cross-References

SkillRelationship
dev-context-engineeringContext quality directly affects AI tool effectiveness — L0-L4 maturity model correlates with metric outcomes
ai-agentsPer-task token economics and agent ROI (this skill covers team/org-level metrics)
qa-observabilityOpenTelemetry integration for automated metric collection
product-managementOKR integration — AI metrics feed into engineering OKRs
startup-business-modelsUnit economics context for ROI calculations
dev-workflow-planningCycle time and planning metrics overlap

Do / Avoid

Do:

  • Start with adoption metrics — you can't optimize what people aren't using
  • Establish baselines *before* rolling out AI tools (8-week minimum)
  • Use the balanced scorecard approach (adoption + velocity + quality + experience)
  • Run quarterly developer experience surveys
  • Report with confidence intervals, not point estimates
  • Cross-reference with context maturity (dev-context-engineering) — structured repos get more AI benefit

Avoid:

  • Don't track individual developer productivity with AI tools
  • Don't use lines of code as a metric for anything
  • Don't measure impact in the first 4 weeks (adoption curve)
  • Don't rely on vendor-published benchmarks without independent validation
  • Don't survey developers without acting on the results
  • Don't compare teams without controlling for confounding variables
  • Don't present ROI without showing the cost model assumptions

Web Verification

55 curated sources in data/sources.json across 7 categories:

CategorySourcesKey Items
Developer Productivity Research~10DORA, SPACE, METR, ETH Zurich, McKinsey, Nicole Forsgren
AI Tool Adoption Data~8GitHub Copilot studies, Stack Overflow, GitClear, Harvard BS
Industry Case Studies~8Block/Square, Stripe, Klarna, Coinbase, Shopify, Amazon
Frameworks & Methodologies~8DX Company, LinearB, Haystack, Jellyfish, Swarmia
Measurement Tools~7Copilot Metrics API, OpenTelemetry, Grafana, PostHog
Consulting Reports~7McKinsey, BCG, HBR, Gartner, Forrester
Academic Research~7arXiv (Peng et al., Ziegler et al., METR), ACM, IEEE

Verify current data before final answers. Priority areas:

  • DORA State of DevOps report updates (annual)
  • GitHub Copilot Metrics API changes
  • New independent productivity studies (academic, not vendor)
  • ETH Zurich context effectiveness research updates
  • METR evaluation methodology updates

Fact-Checking

  • Use web search/web fetch to verify current external facts, versions, pricing, tool features, or published benchmarks before final answers.
  • Prefer independent/academic sources over vendor marketing; report source links and dates.
  • If web access is unavailable, state the limitation and mark guidance as unverified.

Navigation

References

FileContentLines
adoption-metrics.mdAdoption tracking, curve phases, tool-specific data sources, stall patterns~300
productivity-metrics.mdDORA for AI teams, SPACE framework, cycle time decomposition~350
quality-metrics.mdDefect metrics, complexity, test coverage, security, technical debt~280
roi-framework.mdCost/value models, ROI formulas, executive reporting, benchmarks~320
developer-experience-metrics.mdSatisfaction surveys, cognitive load, friction, trust, onboarding~260
benchmarking-methodology.mdA/B comparison, before/after design, statistical rigor, reporting~300

Assets (Copy-Ready Templates)

FilePurpose
metric-dashboard-template.mdThree-tier dashboard layout (Executive / Team Lead / Developer)
adoption-survey-template.md15-question developer survey with Likert scales and scoring
roi-calculator-template.mdSpreadsheet-ready ROI formulas and sensitivity analysis
executive-report-template.mdMonthly 1-page + quarterly deep-dive report templates
experiment-design-template.mdControlled experiment planning with statistical requirements

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Codex

36.57%
按下载量换算120

Claude

29.89%
按下载量换算98

Cursor

20.55%
按下载量换算67

Gemini CLI

10.42%
按下载量换算34

安全审计

Gen Agent Trust Hub

通过

Socket

通过

Snyk

可疑

权限和风险

需要联网

该 Skill 可能需要联网访问来源站点、仓库或外部 API;具体网络访问范围需要结合源码和 README 复核。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。来源安全扫描存在 warning/failed 结果,不能写成本站确认安全。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills