Token导航 LogoToken导航TokenDH.com
研究检索权限需确认github未标认证来源可访问许可证需确认审计未展示

observability-&-monitoring可观察性和监控

Agent Skill

observability-&-monitoring 用于查找、检索和筛选相关信息,适合在 Codex、Claude、Cursor、Gemini CLI 中需要根据关键词、任务场景或来源线索快速定位候选结果时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

12,424

周安装

290

GitHub Stars

公开资料未说明

下载量

3,645
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:observability-&-monitoring(可观察性和监控)
来源仓库:https://github.com/yonatangross/skillforge-claude-plugin
仓库路径:skills/observability-&-monitoring
安装命令:
npx skills add yonatangross/skillforge-claude-plugin --skill "observability-&-monitoring"
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

AgentSkills.tonpx skills
npx skills add yonatangross/skillforge-claude-plugin --skill "observability-&-monitoring"

简介

用于查找、检索和筛选可观察性与监控相关技术方案。

  • 适合在 Codex、Claude、Cursor、Gemini CLI 中根据关键词或运维需求快速定位监控体系。
  • 支持基于来源线索筛选候选结果,可结合仓库路径和原始文档继续核验细节。
  • 安装前需确认权限范围、维护状态及是否会触发日志整理或指标上报操作。
  • observability-&-monitoring 属于研究检索类 Skill,可作为该场景下的辅助能力补充。

SKILL.md

name
monitoring-observability
license
MIT
compatibility
Claude Code 2.1.76+.
description
Monitoring and observability patterns for Prometheus metrics, Grafana dashboards, Langfuse v4 LLM tracing (as_type, score_current_span, should_export_span, LangfuseMedia), and drift detection. Use when adding logging, metrics, distributed tracing, LLM cost tracking, or quality drift monitoring.
tags
[monitoring, observability, prometheus, grafana, langfuse, tracing, metrics, drift-detection, logging]
context
fork
version
3.0.0
author
OrchestKit
user-invocable
false
disable-model-invocation
true
complexity
medium
persuasion-type
reference
targets
version
>=4.0.0
metadata
category
document-asset-creation
allowed-tools
path_patterns
["/metrics/", "/tracing/", "prometheus.*", "grafana/**"]

Monitoring & Observability

Comprehensive patterns for infrastructure monitoring, LLM observability, and quality drift detection. Each category has individual rule files in rules/ loaded on-demand.

Quick Reference

CategoryRulesImpactWhen to Use
Infrastructure Monitoring3CRITICALPrometheus metrics, Grafana dashboards, alerting rules
LLM Observability3HIGHLangfuse tracing, cost tracking, evaluation scoring
Drift Detection3HIGHStatistical drift, quality regression, drift alerting
Silent Failures3HIGHTool skipping, quality degradation, loop/token spike alerting

Total: 12 rules across 4 categories

Quick Start

# Prometheus metrics with RED method
from prometheus_client import Counter, Histogram

http_requests = Counter('http_requests_total', 'Total requests', ['method', 'endpoint', 'status'])
http_duration = Histogram('http_request_duration_seconds', 'Request latency',
    buckets=[0.01, 0.05, 0.1, 0.5, 1, 2, 5])
# Langfuse v4 LLM tracing — semantic as_type + inline scoring
from langfuse import observe, get_client

@observe(as_type="generation", name="analyze_content")
async def analyze_content(content: str):
    get_client().update_current_trace(
        user_id="user_123", session_id="session_abc",
        tags=["production", "orchestkit"],
    )
    result = await llm.generate(content)
    get_client().score_current_span(name="response_quality", value=0.85)
    return result
# PSI drift detection
import numpy as np

psi_score = calculate_psi(baseline_scores, current_scores)
if psi_score >= 0.25:
    alert("Significant quality drift detected!")

Infrastructure Monitoring

Prometheus metrics, Grafana dashboards, and alerting for application health.

RuleFileKey Pattern
Prometheus Metricsrules/monitoring-prometheus.mdRED method, counters, histograms, cardinality
Grafana Dashboardsrules/monitoring-grafana.mdGolden Signals, SLO/SLI, health checks
Alerting Rulesrules/monitoring-alerting.mdSeverity levels, grouping, escalation, fatigue prevention

LLM Observability

Langfuse-based tracing, cost tracking, and evaluation for LLM applications.

RuleFileKey Pattern
Langfuse Tracesrules/llm-langfuse-traces.md@observe decorator, OTEL spans, agent graphs
Cost Trackingrules/llm-cost-tracking.mdToken usage, spend alerts, Metrics API v2
Eval Scoringrules/llm-eval-scoring.mdCustom scores, evaluator tracing, quality monitoring

Drift Detection

Statistical and quality drift detection for production LLM systems.

RuleFileKey Pattern
Statistical Driftrules/drift-statistical.mdPSI, KS test, KL divergence, EWMA
Quality Driftrules/drift-quality.mdScore regression, baseline comparison, canary prompts
Drift Alertingrules/drift-alerting.mdDynamic thresholds, correlation, anti-patterns

Silent Failures

Detection and alerting for silent failures in LLM agents.

RuleFileKey Pattern
Tool Skippingrules/silent-tool-skipping.mdExpected vs actual tool calls, Langfuse traces
Quality Degradationrules/silent-degraded-quality.mdHeuristics + LLM-as-judge, z-score baselines
Silent Alertingrules/silent-alerting.mdLoop detection, token spikes, escalation workflow

Key Decisions

DecisionRecommendationRationale
Metric methodologyRED method (Rate, Errors, Duration)Industry standard, covers essential service health
Log formatStructured JSONMachine-parseable, supports log aggregation
TracingOpenTelemetryVendor-neutral, auto-instrumentation, broad ecosystem
LLM observabilityLangfuse (not LangSmith)Open-source, self-hosted, built-in prompt management
LLM tracing API@observe(as_type=...) + score_current_span()v4: semantic types, inline scoring, span filtering
Langfuse APIsObservations API v2 + Metrics API v2v4 (Mar 2026): faster querying, aggregations at scale
Drift methodPSI for production, KS for small samplesPSI is stable for large datasets, KS more sensitive
Threshold strategyDynamic (95th percentile) over staticReduces alert fatigue, context-aware
Alert severity4 levels (Critical, High, Medium, Low)Clear escalation paths, appropriate response times

Detailed Documentation

ResourceDescription
${CLAUDE_SKILL_DIR}/references/Logging, metrics, tracing, Langfuse, drift analysis guides
${CLAUDE_SKILL_DIR}/checklists/Implementation checklists for monitoring and Langfuse setup
${CLAUDE_SKILL_DIR}/examples/Real-world monitoring dashboard and trace examples
${CLAUDE_SKILL_DIR}/scripts/Templates: Prometheus, OpenTelemetry, health checks, Langfuse

Related Skills

  • defense-in-depth - Layer 8 observability as part of security architecture
  • devops-deployment - Observability integration with CI/CD and Kubernetes
  • resilience-patterns - Monitoring circuit breakers and failure scenarios
  • llm-evaluation - Evaluation patterns that integrate with Langfuse scoring
  • caching - Caching strategies that reduce costs tracked by Langfuse

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

04

需要参考平台分布和安装热度时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Codex

33.21%
按下载量换算1,211

Claude

30.67%
按下载量换算1,118

Cursor

18.66%
按下载量换算680

Gemini CLI

9.3%
按下载量换算339

安全审计

暂无安全审计结果可展示。

权限和风险

权限需确认

当前来源未能明确判断权限范围,默认进入异常复核队列。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills