Token导航 LogoToken导航TokenDH.com
研究检索只读clawhub未标认证来源可访问clear审计通过

vector-databases矢量数据库

Agent Skill

用于辅助数据库表结构、查询语句、迁移脚本和数据维护任务。它适合让 Agent 分析 schema、编写 SQL、排查查询问题、整理索引或生成迁移建议。使用时需要明确数据库类型、连接环境和目标表,区分只读分析与写入变更;涉及删除、更新、迁移和批量导入时,应优先 dry-run、备份或事务保护,避免误操作。

总安装

6,510

周安装

274

GitHub Stars

公开资料未说明

下载量

2,280
OpenClaw

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

MIT-0

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:vector-databases(矢量数据库)
来源仓库:https://github.com/clawkk/vector-databases
安装命令:
openclaw skills install vector-databases
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 OpenClaw 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

ClawHubOpenClaw
openclaw skills install vector-databases

简介

vector-databases 提供矢量数据库的深度操作指南,涵盖嵌入选择与索引算法优化。

  • 适用于需要相似性搜索、混合查询或大规模向量存储的 OpenClaw 项目。
  • 可协助设计召回/延迟权衡策略、过滤条件及成本调整方案,提升检索效率。
  • 安装命令为 openclaw skills install vector-databases,需确认数据库连接与环境权限。
  • 涉及数据写入时应优先 dry-run 或事务保护,避免误删或结构损坏。

SKILL.md

name
vector-databases
description
Deep vector database workflow—embedding choice, index algorithms, recall/latency trade-offs, hybrid search, filtering, operational tuning, and cost. Use when selecting or optimizing Pinecone, Milvus, Qdrant, Weaviate, pgvector, OpenSearch kNN, etc.

Vector Databases (Deep Workflow)

Vector search is approximate nearest neighbor (ANN) at scale—not magic semantic understanding. Success requires embedding model alignment, index parameters, metadata filters, and evaluation against real queries.

When to Offer This Workflow

Trigger conditions:

  • Building RAG, similarity search, dedup, recommendations, anomaly clustering
  • Comparing managed vector DB vs pgvector vs search engine kNN
  • Recall issues, stale vectors, slow queries, or cost explosions

Initial offer:

Use six stages: (1) problem & metrics, (2) embeddings & schema, (3) index & parameters, (4) hybrid & filtering, (5) operations & cost, (6) evaluation & iteration. Confirm scale (vectors, QPS, dimension) and latency SLO.


Stage 1: Problem & Metrics

Goal: Define what “similar” means for the product—not only cosine similarity.

Questions

  1. Query types: short keyword vs long paragraph? multilingual?
  2. Precision vs recall preference: legal/medical may need high precision
  3. Freshness: how often do vectors change? Real-time upserts?

4 Ground truth: any labeled relevant pairs for eval?

Metrics

  • Recall@k, MRR, nDCG when judgments exist; otherwise human spot checks + proxy tasks

Exit condition: Success metric and minimum acceptable recall/latency stated.


Stage 2: Embeddings & Schema

Goal: Stable embedding pipeline with versioning and metadata design.

Embeddings

  • Model choice: domain fit (code vs general text); dimension; distance metric (cosine vs dot vs L2)—match DB defaults
  • Chunking strategy upstream—bad chunks → bad retrieval regardless of DB

Schema

  • Payload/metadata per vector: doc_id, tenant_id, acl, source, timestamps
  • Multi-vector per doc (passages) vs single centroid—trade-offs

Versioning

  • Re-embed all on model change—plan downtime or dual-write period

Exit condition: ID strategy + metadata filter needs documented.


Stage 3: Index & Parameters

Goal: Pick index type and build params for data size and recall.

Common families (vendor-specific names)

  • HNSW: strong latency/recall; memory hungry; tunable efConstruction, M
  • IVF: better memory; needs training nlist; probe tuning
  • PQ/OPQ: compression—recall hit; good for huge scale

Tuning loop

  • Start defaults; sweep parameters with benchmark queries
  • Watch insert throughput during index build on large backfills

Exit condition: Benchmark results: p95 latency vs recall at fixed k.


Stage 4: Hybrid Search & Filtering

Goal: Combine vector similarity with structured constraints—most production needs this.

Patterns

  • Pre-filter metadata (tenant, date) before ANN when supported—verify filter selectivity
  • Hybrid: BM25 + vector with weighted fusion or rerank stage
  • Reranking: cross-encoder on top-k candidates—quality boost, latency cost

Pitfalls

  • Filtering that leaves too few candidates—empty results despite “similar” existing in other tenants

Exit condition: Query plan documented: ANN → filter → rerank (as applicable).


Stage 5: Operations & Cost

Goal: Reliable ingestion, monitoring, and predictable bills.

Ops

  • Upsert idempotency; delete tombstones for compliance
  • Backups, multi-region if needed—eventual consistency semantics per vendor
  • Capacity: memory per node vs sharding; replication factor

Cost

  • Managed per dimension × count; egress; query units—estimate from peak QPS

Exit condition: Runbook for reindex, scaling, and incident “search degraded.”


Stage 6: Evaluation & Iteration

Goal: Continuous improvement with labeled or proxy eval.

Loop

  • Golden query set updated when product changes
  • A/B embedding models or rerankers with guardrails on latency
  • Monitor click-through, thumbs, or human grading in RAG

Debugging bad retrieval

  • Chunk inspection, metadata leaks, wrong tenant filter, stale index

Final Review Checklist

  • [ ] Metrics and embedding/model versioning plan
  • [ ] Index family chosen with benchmark evidence
  • [ ] Hybrid/filter strategy matches product needs
  • [ ] Ops: upsert, delete, scaling, backup understood
  • [ ] Eval set and iteration process in place

Tips for Effective Guidance

  • Never promise “semantic search understands intent”—ground with eval.
  • pgvector vs specialized: trade-offs on scale, ops, features—state honestly.
  • Warn: high-cardinality filters + ANN can be slowdesign metadata carefully.

Handling Deviations

  • Tiny corpus: brute force or simple index may suffice—avoid over-engineering.
  • Multimodal: separate embedding spaces or unified model—fusion strategy required.

适合场景

01

OpenClaw 用户查找和安装 Skill 时

02

用户想查找某类 Agent Skill 时

03

需要根据任务场景推荐可安装能力包时

04

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

能力 5

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

OpenClaw

74.7%
按下载量换算1,703

安全审计

VirusTotal

通过

ClawScan

通过

Static analysis

通过

权限和风险

只读

该 Skill 主要提供规则、说明或参考内容,本身偏只读;真正读写文件、联网或执行命令仍取决于宿主 Agent 的任务。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills