Token导航 LogoToken导航TokenDH.com
研究检索需要联网github未标认证来源可访问clear审计通过

llm-inference-batching-schedulerLLM inference batching scheduler 搜索

Agent Skill

llm-inference-batching-scheduler 用于查找、检索和筛选相关信息,适合在 Codex、Claude、Cursor、Gemini CLI 中需要根据关键词、任务场景或来源线索快速定位候选结果时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

1,023

周安装

41

GitHub Stars

93

下载量

331
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

3

许可证

MIT

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:llm-inference-batching-scheduler(LLM inference batching scheduler 搜索)
来源仓库:https://github.com/letta-ai/skills
仓库路径:skills/llm-inference-batching-scheduler
安装命令:
npx skills add https://github.com/letta-ai/skills --skill llm-inference-batching-scheduler
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。不同来源提供的安装方式可能略有差异;本站展示可直接复制的安装命令,安装前请核对来源页面。

skills.shnpx skills
npx skills add https://github.com/letta-ai/skills --skill llm-inference-batching-scheduler

简介

用于查找、检索和筛选相关信息,适合在 Codex、Claude、Cursor、Gemini CLI 中快速定位候选结果。

  • 支持基于关键词、任务场景或来源线索进行信息提取与整理。
  • 通过 npx skills add 命令从指定 GitHub 仓库安装使用。
  • 安装前需确认权限范围、维护状态及是否触发联网或文件操作。
  • llm-inference-batching-scheduler 属于研究检索类 Skill,可作为该场景下的辅助能力补充。

SKILL.md

LLM Inference Batching Scheduler

This skill provides guidance for solving LLM inference batching and scheduling optimization problems, where requests must be grouped into batches while minimizing cost, padding waste, and latency.

Problem Understanding

Before implementation, thoroughly analyze the problem structure:

Constraint Analysis

  1. Identify all hard constraints - Extract exact limits for:

- Maximum unique shapes allowed (e.g., ≤ 8 shapes across all buckets) - Latency thresholds (P95, P99) - Cost budget thresholds - Padding ratio limits

  1. Compute hard bounds early - Before coding, calculate:

- Minimum possible padding from alignment requirements - Minimum number of batches required for coverage - Maximum achievable efficiency given constraints

  1. Decompose the cost function - Understand each component:

- Per-batch overhead (fixed cost per batch) - Shape compilation costs (often quadratic in sequence length) - Prefill/decode costs (variable per request) - Document as: Cost ≈ overhead × num_batches + shape_compile_cost + prefill_cost + decode_cost

Data Analysis

  1. Profile the request distribution - Examine:

- Distribution of prompt lengths (prompt_len) - Distribution of generation lengths (gen_len) - Identify outliers that may disproportionately impact metrics

  1. Verify coverage requirements - Ensure:

- The largest prompt_len in each bucket is covered by chosen shapes - Edge cases with extreme gen_len values are handled

Implementation Approach

Build Reusable Evaluation Infrastructure

Before iterating on parameters, create a systematic evaluation harness:

1. Write a function that takes parameters (shape_list, gen_bucket_sizes) and returns all metrics
2. Include automatic constraint verification with assertions
3. Enable rapid parameter comparison without manual re-runs

Parameter Search Strategy

Avoid random trial-and-error. Instead:

  1. Grid search for small parameter spaces - When parameters are bounded (e.g., gen_bucket_size in [15, 50]), systematically evaluate combinations
  2. Binary search for single parameters - When optimizing one parameter while holding others fixed, use binary search to find optimal values
  3. Document the optimization landscape - Track which parameter combinations produce which metric values to understand trade-offs

Shape Selection Guidelines

When selecting shapes for sequence length bucketing:

  1. Analyze the length distribution - Choose shapes that minimize padding for the most common lengths
  2. Consider power-of-two or geometric progressions - These often balance coverage vs. shape count
  3. Account for both buckets jointly - If shapes are shared across buckets, optimize globally not independently

Generation Length Bucketing

The gen_len bucketing parameter significantly impacts padding ratio:

  1. Smaller buckets = Lower padding ratio but more batches (higher cost)
  2. Larger buckets = Fewer batches but higher padding from variance in gen_len
  3. Find the sweet spot by computing the padding budget and working backwards

Verification Strategies

Early Constraint Checks

Immediately after generating output, verify:

1. All request IDs appear exactly once
2. Number of unique shapes ≤ limit
3. Each request's prompt_len ≤ assigned shape
4. No missing shapes that would cause alignment failures

Metric Validation

Before considering a solution complete:

  1. Run the official evaluation script (if provided)
  2. Compare all metrics against all thresholds
  3. Check each bucket independently - passing one bucket does not guarantee passing others

Common Verification Failures

Watch for these issues:

  • Missing shapes that cause coverage gaps (e.g., shape 2048 missing when needed)
  • Single-request batches that waste per-batch overhead
  • Shape constraints violated when optimizing buckets independently

Common Pitfalls

Premature Optimization

  • Mistake: Jumping into implementation before understanding mathematical constraints
  • Fix: Spend time upfront computing exact budgets (e.g., "bucket_1 can tolerate at most 25,735 padding tokens")

Insufficient Cost Analysis

  • Mistake: Not understanding which cost component dominates
  • Fix: Compute and document the full cost breakdown before optimizing

Independent Bucket Optimization

  • Mistake: Optimizing each bucket separately when constraints span both
  • Fix: Consider joint optimization, especially for shared shape constraints

Manual Parameter Tuning Loops

  • Mistake: Repeatedly changing parameters manually and re-running
  • Fix: Write a parameter sweep script that tests multiple combinations automatically

Late Discovery of Constraint Violations

  • Mistake: Only checking constraints after extensive iteration
  • Fix: Add assertion checks immediately after output generation

Ignoring Baseline Implementation

  • Mistake: Not analyzing provided baseline code to understand inefficiencies
  • Fix: Study baseline_packer.py (or equivalent) to identify what specifically makes it suboptimal

Optimization Trade-offs

Document these trade-offs explicitly before optimizing:

LeverDecreasesIncreases
More shapesPadding ratioShape compilation cost
Smaller gen bucketsPadding ratioBatch count (cost)
Larger batch sizesPer-batch overheadPotential padding
Tighter shape intervalsPaddingNumber of shapes needed

Recommended Workflow

  1. Analyze - Parse problem, identify constraints, compute hard bounds
  2. Instrument - Build evaluation harness with automatic constraint checking
  3. Baseline - Run baseline solution, understand its deficiencies
  4. Decompose - Break cost into components, identify dominant terms
  5. Search - Use systematic parameter search (grid/binary), not random exploration
  6. Verify - Check all constraints and metrics for all buckets
  7. Iterate - If failing, identify which metric is furthest from threshold and focus there

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

04

需要参考平台分布和安装热度时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

能力 5

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Claude Code

27.44%
按下载量换算91

Gemini CLI

21.13%
按下载量换算70

Codex

17.86%
按下载量换算59

Antigravity

11.72%
按下载量换算39

OpenCode

7.18%
按下载量换算24

windsurf

3.44%
按下载量换算11

安全审计

Gen Agent Trust Hub

通过

Socket

通过

Snyk

通过

权限和风险

需要联网

该 Skill 可能需要联网访问来源站点、仓库或外部 API;具体网络访问范围需要结合源码和 README 复核。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。

来源信息

继续浏览同类 Skills