本文へ移動
cccskills
無料GitHub で公開

vector-search

Semantic vector search with moflo — RAG over your own documents, similarity matching, context-aware retrieval via HNSW (node:sqlite-backed). Use when building retrieval layers for chat, search, or context-assembly.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md5.2 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

MoFlo Vector Search (RAG)

Semantic search over your own documents, backed by moflo's HNSW index in .moflo/moflo.db (node:sqlite, Node 22+ built-in). Small enough to ship in a devDependency; fast enough for interactive retrieval at 100k–1M vectors.

When to Use This vs memory-patterns

  • memory-patterns — structured, namespaced memory you own (sessions, learnings, patterns). Keys matter. Entries are conceptual units.
  • This skill (vector-search) — search over documents you've ingested for retrieval. Entries are content chunks. Keys are just stable IDs for dedupe.

Both use the same index; the difference is how you chunk and what you put in the value field.

Ingest

Two paths. Pick the one that matches your source of truth:

A. Ad-hoc ingest from Claude Code

for (const [id, text] of chunks) {
  await mcp.memory_store({
    namespace: 'docs',        // your RAG corpus
    key: id,                  // stable ID for this chunk (file path + offset, etc.)
    value: text,              // the chunk content — what gets embedded
    tags: ['doc', docType],
    upsert: true,
  });
}

B. Repeatable ingest from the filesystem

Use moflo's shipped indexers (cross-platform, skip unchanged chunks via a hash file):

# Guidance docs → namespace 'guidance'
node .claude/scripts/index-guidance.mjs

# Code structure → namespace 'code-map'
node .claude/scripts/generate-code-map.mjs

# Your own corpus — write a small indexer that loops over files and calls
# memory_store. Model it on bin/index-patterns.mjs.

Retrieve

const hits = await mcp.memory_search({
  namespace: 'docs',
  query: userQuestion,
  limit: 8,
  threshold: 0.35,   // drop irrelevant noise; tune per-corpus
});

// hits[] = [{ key, namespace, value, similarity, ... }]

Plug the retrieved value strings into your prompt as context. Filter or rerank higher up if needed.

Chunking Rules That Actually Matter

  1. One idea per chunk. 200–800 tokens. Smaller for code, larger for prose.
  2. Include enough context for the chunk to stand alone. A bare "it does X" chunk won't retrieve well because the embedding has no subject.
  3. Deduplicate by stable ID. Re-ingest is a no-op if key + value hash match. Cheap to re-run.
  4. Keep metadata minimal in value. The embedding is built from the full value — noisy prefixes dilute the vector. Put structured metadata in tags or a separate meta:${id} key.

Threshold Tuning

The threshold parameter is a similarity floor (0–1, cosine). No universal number.

Workflow for a new corpus:

  1. Pick 10 queries you care about.
  2. Search with threshold: 0.2 and limit: 20 — inspect the ranked results.
  3. Note the similarity score where results go from "relevant" to "noise."
  4. Set threshold just under that cliff. Expect values between 0.3 and 0.6 for most corpora.

Cheaper than trying to predict it.

Hybrid: Semantic + Metadata Filter

moflo's memory_search doesn't have metadata filtering built-in. Do it in two steps:

const candidates = await mcp.memory_search({
  namespace: 'docs',
  query: q,
  limit: 50,           // over-fetch
  threshold: 0.25,
});
const filtered = candidates.filter(c => c.tags?.includes('api-v2'));
const topK = filtered.slice(0, 8);

This is "search then filter" — works well when the filter is permissive. For highly selective filters, reverse the order: list by tag first, then re-rank against the query.

RAG Pattern: Context Assembly

async function answer(question: string) {
  const hits = await mcp.memory_search({
    namespace: 'docs',
    query: question,
    limit: 8,
    threshold: 0.35,
  });

  const context = hits
    .map((h, i) => `[Source ${i + 1}] ${h.value}`)
    .join('\n\n');

  return callModel({
    system: 'Answer using ONLY the sources below. Cite by source number.',
    user: `${context}\n\nQuestion: ${question}`,
  });
}

Anti-Patterns

  • Don't store 50-page documents as single entries. Chunk. Embeddings capture one topic at a time; big values average their meaning into mush.
  • Don't set threshold: 0. You'll get the whole table sorted by similarity — including total noise at the bottom — and waste tokens on worthless context.
  • Don't re-embed on every query. The embedding cost is in the ingest path; queries reuse the HNSW graph. If you find yourself re-embedding, you're doing it wrong.
  • Don't mix namespaces. Keep each corpus in its own namespace so scores are comparable and threshold tuning actually works.

Performance

  • 100k chunks, dim 1536: ~5–15ms per search on commodity hardware.
  • Cold-start: ~5s to load/build HNSW on first query in a process. Keep the process warm.
  • Tuning: see memory-optimization skill for M/ef/quantization knobs.

See Also

  • memory-patterns skill — non-RAG memory usage (sessions, knowledge, patterns)
  • memory-optimization skill — scaling past ~100k vectors
  • bin/semantic-search.mjs / flo-search CLI — shell access to the same index

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

commune

無料

Turn a vague idea into a concrete, actionable spec through a short Socratic dialogue, then hand the result off to an existing moflo surface — a /flo ticket, a spell, or memory. Use BEFORE you have a defined unit of work, when the goal is still fuzzy.

日本語の概要は準備中です。原文の説明を表示しています。

eric-cielo/moflo182026年10月1日 更新

Scaffold new spell step commands and connectors. Use when building new step commands for spells or extending the spell engine with new capabilities. Connectors are for new I/O transport types OR platforms requiring complex multi-step interaction (e.g., browser-based automation).

日本語の概要は準備中です。原文の説明を表示しています。

eric-cielo/moflo182026年10月1日 更新

distill

無料

Alias for /flo-simplify — see that skill's description.

日本語の概要は準備中です。原文の説明を表示しています。

eric-cielo/moflo182026年10月1日 更新

divine

無料

Structured multi-hop web research with explicit confidence gating — plan the inquiry, search (WebSearch/WebFetch), score your own confidence, and keep digging until the answer is well-supported or a hop cap is hit, then emit a cited synthesis. Learns across sessions by storing each research case to memory and reusing prior strategies. Use when a question needs more than one search — comparisons, current-best-practice questions, anything where a single lookup leaves you unsure.

日本語の概要は準備中です。原文の説明を表示しています。

eric-cielo/moflo182026年10月1日 更新

eldar

無料

Consult the Eldar — audit a project's moflo + Claude Code setup for portable, high-leverage gaps and guide remediation. Default mode is read-only audit with severity-ranked findings; --fix presents an interactive triage menu and walks the user through each chosen fix (healer, missing CLAUDE.md, sparse guidance, hook/MCP wiring, empty memory namespaces, stack→guidance gaps). Use when starting in a new project, when Claude feels lost or inefficient, when guidance/CLAUDE.md is sparse, or as a periodic health check.

日本語の概要は準備中です。原文の説明を表示しています。

eric-cielo/moflo182026年10月1日 更新

flfl

無料

Run /fl on a ticket with moflo's three standing considerations loaded first — cross-platform (Rule

日本語の概要は準備中です。原文の説明を表示しています。

eric-cielo/moflo182026年10月1日 更新

eric-cielo のスキルをすべて見る

このスキルの問題を報告する