Token导航 LogoToken导航TokenDH.com
研究检索需要联网github未标认证来源可访问许可证需确认审计未展示

ux-metrics-%26-measurement用户体验指标 %26 测量

Agent Skill

用于辅助界面设计、视觉规范、排版、配色、布局和交互体验优化。它适合让 Agent 根据产品场景整理页面结构、生成 UI 方案、检查视觉一致性或改进组件层级。使用时需要结合现有品牌、设计系统和用户任务,不应只堆装饰元素;涉及真实页面改动时,应通过截图或浏览器预览检查文本溢出、对齐和响应式表现。

总安装

23,112

周安装

839

GitHub Stars

3

下载量

9,230
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:ux-metrics-%26-measurement(用户体验指标 %26 测量)
来源仓库:https://github.com/phazurlabs/ux-ui-mastery
仓库路径:skills/ux-metrics-%26-measurement
安装命令:
npx skills add https://github.com/phazurlabs/ux-ui-mastery --skill 'UX Metrics & Measurement'
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

skills.shnpx skills
npx skills add https://github.com/phazurlabs/ux-ui-mastery --skill 'UX Metrics & Measurement'

简介

用于评估用户体验效果并提供量化依据。

  • 适合选择合适指标并建立测量体系。ux-metrics-%26-measurement 属于研究检索类 Skill,可作为该场景下的辅助能力补充。
  • 使用时需区分行为指标与满意度指标的应用场景。
  • 帮助定位问题根源并衡量优化成效。
  • 安装方式:github,支持多平台技能扩展。

SKILL.md

UX Metrics & Measurement — Quantifying User Experience

Measurement Philosophy

Measure what matters, not what is easy. Vanity metrics (page views, downloads, time on site) tell you what happened but not whether the experience was good. Meaningful UX metrics connect user behavior to user satisfaction to business outcomes. Every metric in the system should answer one of three questions: Is it usable? Is it useful? Is it desirable?

The Measurement Hierarchy

  1. Behavioral metrics (what users do) are more reliable than attitudinal metrics (what users say).
  2. Task-level metrics (can users accomplish goals?) matter more than session-level metrics (how long did users stay?).
  3. Leading indicators (early signals of success/failure) are more actionable than lagging indicators (outcomes measured after the fact).
  4. Benchmarked metrics (compared to industry or historical baseline) are more meaningful than absolute metrics (numbers without context).

HEART Framework (Google)

The HEART framework provides a structured approach to selecting user-centered metrics across five dimensions. Developed at Google, it maps product goals to measurable signals to concrete metrics.

Five Dimensions

Happiness — Subjective user satisfaction and sentiment

  • Signals: survey responses, ratings, sentiment in feedback
  • Example metrics: satisfaction score (CSAT), NPS, SUS score, ease-of-use rating
  • Measurement: post-task surveys, in-app ratings, periodic satisfaction surveys

Engagement — Depth and frequency of user interaction

  • Signals: session frequency, feature usage, content consumption
  • Example metrics: DAU/MAU ratio, actions per session, feature adoption rate, session depth
  • Measurement: product analytics, event tracking

Adoption — New users successfully onboarding

  • Signals: signups, onboarding completion, first key action
  • Example metrics: signup-to-activation rate, onboarding completion rate, time to first value
  • Measurement: funnel analytics, cohort analysis

Retention — Users continuing to derive value over time

  • Signals: return visits, subscription renewals, continued usage
  • Example metrics: D1/D7/D30 retention, churn rate, renewal rate, feature retention
  • Measurement: cohort retention curves, survival analysis

Task Success — Ability to complete intended goals effectively and efficiently

  • Signals: task completion, errors, time spent, help requests
  • Example metrics: task success rate, time on task, error rate, lostness score
  • Measurement: usability testing, analytics funnels, error logging

Goals-Signals-Metrics Process

For each product area, work through three columns:

  1. Goal: What is the desired user outcome? ("Users can find relevant products quickly")
  2. Signal: What user behavior indicates progress toward the goal? ("Users complete searches and click results")
  3. Metric: How do you quantify the signal? ("Search-to-click rate, average time from search to product page view")

This process prevents metric selection from being arbitrary. Every metric traces back to a user goal.

System Usability Scale (SUS)

The most widely used standardized usability questionnaire. Ten questions on a five-point Likert scale, producing a score from 0-100.

Scoring and Interpretation

  • Raw score calculation: alternate positive/negative items, subtract/add from scale anchors, multiply by 2.5.
  • Average SUS score across studies: 68 (this is the 50th percentile, not a passing grade).
  • Adjective scale mapping: <50 = Awful, 50-67 = Poor to OK, 68 = Passable, 68-80 = Good, 80-90 = Excellent, 90+ = Best Imaginable.
  • Grade scale: <60 = F, 60-69 = D, 70-79 = C, 80-89 = B, 90+ = A.
  • Percentile benchmarks: 68 = 50th, 74 = 64th, 80 = 85th, 86 = 96th.

When to Use SUS

  • After usability testing sessions (post-study).
  • Baseline measurement before redesign.
  • Comparative evaluation between design alternatives.
  • Longitudinal tracking across releases.
  • Industry benchmarking.

Limitations

  • SUS measures perceived usability, not actual usability. A user may rate a system highly despite making errors.
  • The 10-item format is not ideal for quick pulse checks — use UMUX-Lite (2 items) for lightweight measurement.
  • SUS does not diagnose specific issues — it is a thermometer, not a diagnostic tool.

UEQ (User Experience Questionnaire)

The UEQ measures six dimensions of user experience: Attractiveness, Perspicuity (clarity), Efficiency, Dependability, Stimulation, and Novelty. It uses 26 semantic differential items (word pairs).

When to Use UEQ Over SUS

  • When you need dimensional insight beyond a single usability score.
  • When hedonic qualities (stimulation, novelty) matter alongside pragmatic qualities (perspicuity, efficiency, dependability).
  • When comparing products across multiple experience dimensions.
  • UEQ+ (modular version) allows selecting only relevant scales.

UMUX-Lite

A two-item questionnaire that correlates strongly with SUS (r = 0.83). Useful for: in-app pulse surveys, post-task quick checks, high-frequency measurement where survey fatigue is a concern.

Items: (1) "[System] capabilities meet my requirements." (2) "[System] is easy to use." Both rated on a seven-point Likert scale.

SUPR-Q (Standardized User Experience Percentile Rank Questionnaire)

Web-specific measurement covering four factors: Usability, Trust/Credibility, Appearance, and Loyalty. Produces a percentile rank against a database of 200+ websites. Best for benchmarking web experiences against industry competitors.

Task-Based Metrics

Task Success Rate

The most fundamental UX metric. Percentage of users who successfully complete a defined task.

  • Binary success: Did the user complete the task? Yes/No.
  • Partial success: Weighted scoring for partially completed tasks (e.g., 0 = fail, 0.5 = completed with significant difficulty, 1 = success).
  • Benchmark: Industry average for web tasks is approximately 78%. Below 70% indicates serious usability problems.

Time on Task

How long users take to complete a task. Lower is generally better, but context matters — rushed completion may indicate skipped steps.

  • Measure from task start to task completion.
  • Use geometric mean (not arithmetic mean) because time data is typically right-skewed.
  • Compare against baseline or competitor benchmarks.
  • Distinguish between first-time and repeat users.

Error Rate

Frequency and severity of errors during task completion.

  • Error frequency: Number of errors per task attempt.
  • Error taxonomy: Classify errors as slips (wrong action, right intention), mistakes (wrong intention), or system errors (not user-caused).
  • Recovery rate: Percentage of errors the user successfully recovers from.
  • Error impact: Which errors cause task abandonment versus temporary friction.

Lostness Score

Measures navigation efficiency. Calculated as: Lostness = sqrt((N/S - 1)^2 + (R/N - 1)^2) where N = number of pages visited, S = minimum pages needed, R = unique pages visited.

  • Score of 0 = optimal path. Score above 0.4 = user is seriously lost.
  • Useful for evaluating information architecture and navigation design.

Behavioral Analytics

Funnel Analysis

Track user progression through multi-step flows. Identify where users drop off and why.

  • Define funnels for every critical flow: onboarding, checkout, feature activation, upgrade.
  • Measure step-to-step conversion rate, overall completion rate, and time per step.
  • Segment funnels by user cohort, acquisition channel, device, and user experience level.
  • Investigate drop-off points with session recordings and qualitative research.

Cohort Analysis

Group users by shared characteristic (signup date, acquisition channel, plan tier) and track behavior over time.

  • Retention cohorts: What percentage of users who signed up in Week 1 are still active in Week 4, 8, 12?
  • Feature cohorts: How does behavior differ between users who adopted Feature X versus those who did not?
  • Identify the "aha moment" — the action most correlated with long-term retention.

Retention Curves

Plot the percentage of active users over time since signup.

  • Flattening curve: Healthy product — a stable base of users finds lasting value.
  • Declining curve: Problem — users are not finding enough value to stay.
  • Smiling curve: Recovery — initial drop-off followed by increasing engagement (often from re-engagement campaigns).

Experimentation

A/B Testing

Compare two design variants with random user assignment to determine which performs better on a defined metric.

  • Hypothesis template: "If we [change], then [metric] will [improve/decrease] by [amount] because [reasoning]."
  • Sample size: Calculate before launching using a power calculator. Specify minimum detectable effect, significance level (typically 0.05), and power (typically 0.80).
  • Duration: Run for at least one full business cycle (minimum 1-2 weeks). Do not peek at results and stop early.
  • One variable: Test one change at a time for clear causation. Multivariate testing requires exponentially larger samples.
  • Guard rails: Monitor guardrail metrics (crash rate, load time, error rate) to ensure the variant does not cause harm even if the primary metric improves.

Sequential Testing

Alternative to fixed-horizon A/B testing. Allows checking results continuously without inflating false positive rate. Useful when you need faster decisions or have variable traffic.

Benchmarking

Industry Benchmarks

  • SUS: Average across studies is 68. B2B software averages 65-72. Consumer apps average 70-78.
  • Task success rate: Web average 78%. E-commerce checkout 65-70%. Mobile forms 60-75%.
  • NPS: Average varies by industry. SaaS averages 30-40. B2C averages 40-60.
  • Time on task: Highly variable. Compare against your own historical baseline rather than cross-industry.

Historical Tracking

  • Track key metrics across every release. Plot trendlines.
  • Set regression thresholds: alert when a metric drops more than 5% from baseline.
  • Celebrate improvements with the team — measurement should drive positive reinforcement.

AI-Specific UX Metrics

  • Trust calibration score: How closely does user confidence in AI output match actual AI accuracy?
  • Appropriate reliance rate: Percentage of times users correctly accept accurate AI outputs and correctly reject inaccurate ones.
  • AI feature adoption: Opt-in rate, prompt frequency, feature retention over time.
  • Correction frequency: How often users edit or override AI outputs.
  • Time saved: Measured difference in task time with AI versus without.

Design System Metrics

  • Adoption rate: Percentage of UI built with design system components versus custom implementations.
  • Component reuse: Average instances per component across products.
  • Consistency score: Visual regression pass rate across products.
  • Contribution rate: External PRs per month to the design system.
  • Developer satisfaction: Quarterly NPS survey of design system consumers.
  • Design system ROI: [(Time efficiency gains + Quality gains + Scale gains + Consistency gains) - (Initial cost + Maintenance cost)] / (Initial + Maintenance) x 100%.
  • Sparkbox research data: 38% design efficiency, 31% dev efficiency, 228% higher ROI with mature DesignOps.

Cross-Referencing

  • For heuristic evaluation methodology, reference nng-ux-heuristics.
  • For research methodology, reference ux-research-methods.
  • For AI-specific metrics detail, reference ux-metrics-measurement/references/ai-ux-metrics-experimentation.
  • For design system metrics detail, reference design-systems-architecture/references/governance-scaling.
  • For usability testing protocols that produce metrics, reference ux-research-methods/references/research-protocols.

v3.0 Cross-References

The v3.0 upgrade adds reference materials that extend measurement into design system maturity, AI-specific quality metrics, and notification effectiveness.

Design System Maturity Metrics See design-systems-architecture/references/maturity-model-multi-brand.md for the 5-level design system maturity assessment rubric with quantifiable metrics at each level — including adoption rate thresholds (Level 3 requires 60%+ component coverage), contribution velocity benchmarks, token coverage ratios, cross-platform parity scores, and governance process maturity indicators. This reference provides the structured assessment framework that complements the Design System Metrics section above, enabling teams to measure where they stand and define concrete advancement targets.

AI-Specific Quality and Trust Metrics See agentic-ai-generative-ux/references/llm-hallucination-design-guardrails.md for expanded AI-specific metrics beyond the AI-Specific UX Metrics section above — including hallucination rate measurement methodology (per-claim factual accuracy scoring), confidence calibration metrics (Expected Calibration Error measuring alignment between stated confidence and actual correctness), trust calibration accuracy (user trust vs. system reliability correlation), verification engagement rate (how often users check AI citations), and AI quality gate pass/fail rates for production deployment monitoring. These metrics are essential for any team shipping LLM-powered features.

Notification Effectiveness Metrics See performance-states-patterns/references/notification-system-design.md for metrics specific to notification system evaluation — including notification-to-action rate (percentage of notifications that lead to meaningful user action), dismissal rate by notification type, opt-out rate trending, notification fatigue indicators (declining engagement over time), time-to-action measurement, and notification channel effectiveness comparison (push vs. in-app vs. email). These metrics connect directly to the HEART framework's Engagement dimension for notification-driven features.

Key Sources

  • Google HEART Framework (Rodden, Hutchinson, Fu)
  • Sauro, J. & Lewis, J.: "Quantifying the User Experience"
  • MeasuringU: 48 UX Metrics (2025)
  • Brooke, J.: SUS — A Quick and Dirty Usability Scale
  • Schrepp, M.: User Experience Questionnaire (UEQ)
  • Frontiers in Computer Science: UX Measurement Frameworks
  • Sparkbox: Design System ROI Research
  • NNG Group: UX Maturity Model, DesignOps 101

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Codex

37.15%
按下载量换算3,429

Claude

27.08%
按下载量换算2,499

Cursor

18.32%
按下载量换算1,691

Gemini CLI

9.97%
按下载量换算920

安全审计

暂无安全审计结果可展示。

权限和风险

需要联网

该 Skill 可能需要联网访问来源站点、仓库或外部 API;具体网络访问范围需要结合源码和 README 复核。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills