Token导航 LogoToken导航TokenDH.com
开发操作浏览器clawhub未标认证来源可访问clear审计通过

scrapling-extract刮痧提取物

Agent Skill

scrapling-extract 用于处理浏览器自动化、网页检查和页面信息提取,适合在 OpenClaw 中需要让 Agent 打开页面、读取网页或验证前端流程时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

19,183

周安装

776

GitHub Stars

1

下载量

6,022
OpenClaw

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

MIT-0

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:scrapling-extract(刮痧提取物)
来源仓库:https://github.com/piyushzinc/scrapling-extract
安装命令:
openclaw skills install scrapling-extract
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 OpenClaw 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

ClawHubOpenClaw
openclaw skills install scrapling-extract

简介

scrapling-extract 基于 Python Scrapling 库,支持静态与 JS 渲染页面抓取。

  • 适用于 OpenClaw 中提取表格、列表或评论区等结构化内容。
  • 可选择无头浏览模式绕过验证码与反爬策略。
  • 需确认是否支持自定义请求头与代理配置,并评估网络延迟影响。
  • 适用于教育资料抓取或开源项目文档索引。

SKILL.md

name
scrapling
description
Web scraping and data extraction using the Python Scrapling library. Use to scrape static HTML pages, JavaScript-rendered pages (Playwright), and anti-bot or Cloudflare-protected sites (stealth browser). Supports CSS selectors, XPath, adaptive DOM relocation so selectors survive site redesigns, session-based scraping with cookie persistence, and outputs to JSON or Markdown. Use when asked to scrape a URL, extract text/links/tables/prices from a webpage, crawl a site, or automate web data collection.

Scrapling

Extract structured website data with resilient selection patterns, adaptive relocation, and the right Scrapling fetcher mode for each target.

Workflow

  1. Identify target type before writing code:

- Use Fetcher for static pages and API-like HTML responses. - Use DynamicFetcher when JavaScript rendering is required. - Use StealthyFetcher when anti-bot protection or browser fingerprinting issues are likely.

  1. Choose output contract first:

- Return JSON for pipelines/automation. - Return Markdown/text for summarization or RAG ingestion. - Keep stable field names even if selector strategy changes.

  1. Implement selectors in this order:

- Start with CSS selectors and pseudo-elements (for example ::text, ::attr(href)). - Fall back to XPath for ambiguous DOM structure. - Enable adaptive relocation for brittle or changing pages.

  1. Add safety controls:

- Respect target site terms and legal boundaries. - Add timeouts, retries, and explicit error handling. - Log status code, URL, and selector misses for debugging.

  1. Validate on at least 2 pages:

- Test one happy path and one edge case page. - Confirm required fields are non-empty. - Keep extraction deterministic (no hidden random choices).

Quick Setup

  1. Install base package:

- pip install scrapling

  1. Install fetchers when browser-based fetching is needed:

- pip install "scrapling[fetchers]" - scrapling install - python3 -m playwright install (required for DynamicFetcher and StealthyFetcher)

  1. Install optional extras as needed:

- pip install "scrapling[shell]" for shell + extract commands - pip install "scrapling[ai]" for MCP capabilities

Execution Patterns

Pattern: One-off terminal extraction

Use Scrapling CLI for fastest no-code extraction:

scrapling extract get "https://example.com" content.md --css-selector "main"

Pattern: Python extraction script

Use the bundled helper:

# Static page (default)
python scripts/extract_with_scrapling.py --url "https://example.com" --css "h1::text"

# JavaScript-rendered page
python scripts/extract_with_scrapling.py --url "https://example.com" --fetcher dynamic --css "h1::text"

# Anti-bot protected page
python scripts/extract_with_scrapling.py --url "https://example.com" --fetcher stealthy --css "h1::text"

Pattern: Session-based scraping

Use session classes when cookies/state must persist across requests.

from scrapling.fetchers import FetcherSession

session = FetcherSession()
login_page = session.post("https://example.com/login", data={"user": "...", "pass": "..."})
protected_page = session.get("https://example.com/dashboard")
headline = protected_page.css_first("h1::text")

Use StealthySession or DynamicSession as drop-in replacements for anti-bot or JS-rendered targets.

Pattern: DOM change resilience

Use auto_save=True on initial capture and retry with adaptive selection on later runs when selectors break.

from scrapling.fetchers import Fetcher

# First run: saves DOM snapshot so adaptive relocation can work later
page = Fetcher.auto_match("https://example.com", auto_save=True, disable_adaptive=False)
price = page.css_first(".price::text")

# Later runs: automatically relocates the selector even if the DOM changed
page = Fetcher.auto_match("https://example.com", auto_save=False, disable_adaptive=False)
price = page.css_first(".price::text")

References

适合场景

01

OpenClaw 用户查找和安装 Skill 时

02

用户想查找某类 Agent Skill 时

03

需要根据任务场景推荐可安装能力包时

04

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

能力 5

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

OpenClaw

74.11%
按下载量换算4,463

安全审计

VirusTotal

通过

ClawScan

通过

Static analysis

未展示

权限和风险

操作浏览器

该 Skill 可能涉及浏览器控制能力,使用时可能读取或操作网页内容,需要在受控环境中确认权限边界。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills