Token导航 LogoToken导航TokenDH.com
研究检索需要联网clawhub未标认证来源可访问clear审计提醒

bio-ontology-mapper生物本体映射器

Agent Skill

bio-ontology-mapper 用于查找、检索和筛选相关信息,适合在 OpenClaw 中需要根据关键词、任务场景或来源线索快速定位候选结果时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

9,032

周安装

384

GitHub Stars

公开资料未说明

下载量

3,164
OpenClaw

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

MIT-0

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:bio-ontology-mapper(生物本体映射器)
来源仓库:https://github.com/renhaosu2024/bio-ontology-mapper
安装命令:
openclaw skills install bio-ontology-mapper
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 OpenClaw 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

ClawHubOpenClaw
openclaw skills install bio-ontology-mapper

简介

bio-ontology-mapper实现生物医学文本到标准本体的映射转换。

  • 主要服务于医学研究和临床数据标准化处理场景。
  • 支持SNOMED CT、MeSH、ICD-10等标准术语体系。
  • 使用前需确认文本类型和目标本体系统的匹配性。
  • 建议了解具体的术语映射规则和数据处理流程。bio-ontology-mapper 属于研究检索类 Skill,可作为该场景下的辅助能力补充。

SKILL.md

name
bio-ontology-mapper
description
Map unstructured biomedical text to standardized ontologies (SNOMED CT,
allowed-tools
[Read, Write, Bash, Edit]
license
MIT
metadata
skill-author
AIPOCH

Bio-Ontology Mapper

Overview

Biomedical terminology normalization tool that maps free-text clinical and scientific concepts to standardized ontologies for semantic interoperability and data harmonization.

Key Capabilities:

  • Multi-Ontology Support: SNOMED CT, MeSH, ICD-10, LOINC, RxNorm
  • Entity Extraction: NER for diseases, symptoms, procedures, drugs
  • Fuzzy Matching: Handle typos, abbreviations, and synonyms
  • Confidence Scoring: Reliability metrics for each mapping
  • Batch Processing: Normalize large datasets efficiently
  • Cross-Mapping: Translate between ontology systems

When to Use

✅ Use this skill when:

  • Normalizing clinical notes for EHR integration
  • Standardizing terminology for multi-site studies
  • Mapping legacy data to modern ontologies
  • Preparing data for clinical data warehouses
  • Converting free-text to coded data for analysis
  • Building semantic search for biomedical literature
  • Teaching biomedical informatics principles

❌ Do NOT use when:

  • Clinical diagnosis or decision support → Use clinical decision tools
  • Real-time patient care → Latency too high for acute settings
  • Replacing expert coding → Use for pre-coding, final review needed
  • Processing PHI without de-identification → Ensure HIPAA compliance

Integration:

  • Upstream: clinical-data-cleaner (data preparation), ehr-semantic-compressor (text extraction)
  • Downstream: clinical-data-cleaner (SDTM mapping), unstructured-medical-text-miner (NLP pipelines)

Core Capabilities

1. Entity Recognition and Mapping

Extract and map biomedical entities to ontologies:

from scripts.mapper import BioOntologyMapper

mapper = BioOntologyMapper()

# Map clinical text
result = mapper.map_text(
    text="Patient has diabetes and hypertension, taking metformin",
    ontologies=["snomed", "mesh", "rxnorm"],
    confidence_threshold=0.7
)

for entity in result.entities:
    print(f"{entity.text} → {entity.concept_id} ({entity.ontology})")
    print(f"  Preferred: {entity.preferred_term}")
    print(f"  Confidence: {entity.confidence:.2f}")

Supported Ontologies:

OntologyDomainUse Case
SNOMED CTClinicalEHR interoperability
MeSHLiteraturePubMed indexing
ICD-10BillingDiagnosis codes
LOINCLabsTest result standardization
RxNormDrugsMedication normalization
HGNCGenesGene name standardization

2. Cross-Ontology Translation

Map concepts between different ontologies:

# Cross-map SNOMED to ICD-10
translation = mapper.cross_map(
    source_id="22298006",  # SNOMED: Myocardial infarction
    source_ontology="snomed",
    target_ontology="icd10"
)

print(f"ICD-10: {translation.target_id} - {translation.target_term}")
# Output: I21.9 - Acute myocardial infarction, unspecified

Cross-Mapping Coverage:

  • SNOMED CT ↔ ICD-10-CM (clinical modifications)
  • MeSH ↔ SNOMED CT (literature to clinical)
  • RxNorm ↔ ATC (drug classifications)
  • LOINC ↔ SNOMED (lab to clinical)

3. Batch Normalization

Process large datasets:

# Batch process CSV
results = mapper.batch_map(
    input_file="clinical_terms.csv",
    text_column="diagnosis_description",
    ontologies=["snomed", "icd10"],
    output_format="csv",
    max_workers=4
)

# Results include:
# - Original term
# - Mapped concept ID
# - Confidence score
# - Alternative mappings (if ambiguous)

Performance:

  • ~100 terms/second (with caching)
  • ~20 terms/second (API lookup)
  • Parallel processing for large datasets

4. Confidence Scoring and Validation

Assess mapping reliability:

scoring = mapper.score_mapping(
    term="heart attack",
    candidate="22298006",  # Myocardial infarction
    factors=["string_similarity", "context_match", "frequency"]
)

print(f"Overall confidence: {scoring.confidence:.2f}")
print(f"Breakdown: {scoring.factors}")

Scoring Factors:

  • String similarity: Levenshtein distance, n-grams
  • Context match: Surrounding words alignment
  • Frequency: Common usage in corpus
  • Semantic similarity: Vector embeddings

Common Patterns

Pattern 1: Clinical Note Normalization

Scenario: Convert free-text diagnoses to SNOMED codes.

# Normalize clinical notes
python scripts/main.py \
  --input notes.csv \
  --column diagnosis_text \
  --ontology snomed \
  --threshold 0.8 \
  --output coded_diagnoses.csv

# Results: "heart attack" → 22298006 (Myocardial infarction)

Post-Processing:

  • Review low-confidence mappings (<0.8)
  • Handle ambiguous terms manually
  • Validate against clinical context

Pattern 2: Literature Indexing

Scenario: Map research paper keywords to MeSH.

# Map keywords to MeSH
mesh_terms = mapper.map_to_mesh(
    keywords=["cancer immunotherapy", "checkpoint inhibitors", "PD-1"],
    include_tree_numbers=True,
    include_qualifiers=True
)

for term in mesh_terms:
    print(f"{term.input} → {term.descriptor}")
    print(f"  Tree: {term.tree_numbers}")
    print(f"  Entry terms: {term.synonyms}")

Pattern 3: Drug Name Normalization

Scenario: Standardize medication names across datasets.

# Normalize drug names
drugs = ["Tylenol", "Advil", "Motrin", "acetaminophen"]

for drug in drugs:
    result = mapper.map_to_rxnorm(drug)
    print(f"{drug} → {result.rxcui}: {result.name}")
    # Tylenol → 161: Acetaminophen
    # Advil → 5640: Ibuprofen
    # Motrin → 5640: Ibuprofen

Pattern 4: EHR Data Harmonization

Scenario: Merge data from multiple hospital systems.

# Harmonize diagnoses from 3 hospitals
python scripts/main.py \
  --batch \
  --inputs "hospital_a.csv,hospital_b.csv,hospital_c.csv" \
  --target-ontology snomed \
  --cross-map-to icd10 \
  --output harmonized_data.csv

Complete Workflow Example

From free-text to coded database:

from scripts.mapper import BioOntologyMapper
from scripts.validator import MappingValidator

# Initialize
mapper = BioOntologyMapper()
validator = MappingValidator()

# Step 1: Extract entities from text
clinical_note = "Patient has Type 2 diabetes and hypertension..."
entities = mapper.extract_entities(clinical_note)

# Step 2: Map to SNOMED
mappings = []
for entity in entities:
    mapping = mapper.map_to_snomed(
        entity.text,
        context=clinical_note,
        top_n=3
    )
    mappings.append(mapping)

# Step 3: Validate mappings
for mapping in mappings:
    validation = validator.validate(
        mapping,
        check_clinical_plausibility=True
    )
    if not validation.is_valid:
        print(f"Review needed: {mapping}")

# Step 4: Export to database format
db_records = [m.to_database_record() for m in mappings]

Quality Checklist

Pre-Mapping:

  • [ ] Text preprocessed (lowercase, punctuation handled)
  • [ ] Abbreviations expanded where possible
  • [ ] Language identified (multilingual support)

During Mapping:

  • [ ] Confidence threshold appropriate (>0.7 for clinical)
  • [ ] Multiple candidates considered for ambiguous terms
  • [ ] Context used for disambiguation

Post-Mapping:

  • [ ] Low-confidence mappings flagged for review
  • [ ] Unmapped terms logged
  • [ ] CRITICAL: Clinical expert validation for high-stakes use

Before Production:

  • [ ] Mapping accuracy validated on gold standard
  • [ ] False positive rate acceptable (<5%)
  • [ ] Recall acceptable for use case (>90%)
  • [ ] API rate limits respected

Common Pitfalls

Mapping Errors:

  • Abbreviation ambiguity → "MI" = Myocardial infarction OR Michigan

- ✅ Use context; flag for manual review

  • Outdated terms → Old terminology not in current ontology

- ✅ Use historical mappings; update terminology

  • False confidence → High score for wrong concept

- ✅ Always review top-3 candidates

Technical Issues:

  • API failures → No local fallback

- ✅ Implement caching; use local reference files

  • Version mismatches → Different ontology versions

- ✅ Track ontology version used

  • PHI exposure → Sending patient data to external APIs

- ✅ De-identify before API calls; use local processing when possible

References

Available in references/ directory:

  • snomed_ct_guide.md - SNOMED CT hierarchy and relationships
  • mesh_structure.md - MeSH tree structure and qualifiers
  • ontology_mappings.md - Crosswalks between systems
  • nlp_best_practices.md - Biomedical text processing
  • api_documentation.md - External service integration
  • validation_datasets.md - Gold standard test sets

Scripts

Located in scripts/ directory:

  • main.py - CLI interface for mapping
  • mapper.py - Core ontology mapping engine
  • extractor.py - Named entity recognition
  • cross_mapper.py - Ontology-to-ontology translation
  • scorer.py - Confidence calculation
  • batch_processor.py - Large dataset handling
  • validator.py - Mapping quality checks
  • caching.py - Local storage for frequent lookups

Limitations

  • Ambiguity: Many-to-many mappings common; context required
  • Coverage: Rare diseases and new concepts may not be in ontologies
  • Versioning: Ontology updates can change mappings over time
  • Language: Best support for English; other languages limited
  • Real-time: Not suitable for time-critical clinical applications
  • API Dependency: Requires internet for most lookups (caching helps)

⚠️ Critical: Ontology mapping is for research and data integration, not clinical decision-making. Always validate mappings with domain experts before use in patient care contexts. Never process PHI without appropriate de-identification and compliance measures.

Parameters

ParameterTypeDefaultDescription
--termstrRequiredSingle term to map
--inputstrRequiredInput file path
--outputstrRequiredOutput file path
--ontologystr'both'
--thresholdfloat0.7
--formatstr'json'
--use-apistrRequiredUse UMLS/MeSH APIs
--api-keystrRequired

适合场景

01

OpenClaw 用户查找和安装 Skill 时

02

用户想查找某类 Agent Skill 时

03

需要根据任务场景推荐可安装能力包时

04

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

能力 5

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

OpenClaw

81.67%
按下载量换算2,584

安全审计

VirusTotal

通过

ClawScan

可疑

Static analysis

通过

权限和风险

需要联网

该 Skill 可能需要联网访问来源站点、仓库或外部 API;具体网络访问范围需要结合源码和 README 复核。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。来源安全扫描存在 warning/failed 结果,不能写成本站确认安全。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills