本文へ移動
cccskills
無料GitHub で公開

memory-optimization

Tune moflo's memory stack for speed, RAM, and index quality. Covers HNSW parameters (M, efConstruction, ef), vector quantization, batch operations, and common bottlenecks. Use when scaling past ~100k entries or when search latency regresses.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md4.9 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

MoFlo Memory Optimization

When the default moflo memory settings stop being enough — past ~100k entries, or when p95 search latency climbs — these are the levers.

HNSW Parameters

HNSW has three knobs. They trade build time, query time, memory, and recall.

import { HNSWIndex } from 'moflo/dist/src/cli/memory/index.js';

const index = new HNSWIndex({
  dimensions: 1536,        // must match your embedding model
  maxElements: 1_000_000,  // pre-allocated capacity
  M: 16,                   // graph connectivity (default 16)
  efConstruction: 200,     // build-time search width (default 200)
  metric: 'cosine',        // 'cosine' | 'l2' | 'ip'
});
KnobHigherLowerWhen to change
Mbetter recall, more RAM (~2×M pointers per point)less RAM, worse recallBump to 32–64 if recall@10 < 0.95; drop to 8 if memory-bound
efConstructionbetter index quality, slower buildfaster build, worse queries200–400 is sweet spot; only lower in test fixtures
ef (search-time, passed to search())better recall, slower queriesfaster queries, worse recallStart at 2×k, raise until recall plateaus

Rule of thumb: M and efConstruction are set once. ef is the runtime dial.

Quantization

moflo memory supports scalar quantization (Float32 → Int8) for a ~4× memory reduction with a ~1-2% recall hit. Turn it on when the index doesn't fit comfortably in RAM.

const index = new HNSWIndex({
  dimensions: 1536,
  maxElements: 5_000_000,
  quantization: {
    enabled: true,
    type: 'scalar',   // scalar (Int8) is the supported path
    rebuildThreshold: 10_000,
  },
});

Measure recall before/after on your own query distribution — public benchmarks don't predict your domain.

Batch Operations

Single-entry writes pay the HNSW insert cost per call. For bulk ingest, batch:

const entries: Array<[string, Float32Array]> = buildCorpus();

// Parallelise at the adapter level; don't await sequentially.
await Promise.all(
  entries.map(([id, vec]) => index.addPoint(id, vec))
);

// Or via MCP for moflo-native batch into .swarm/memory.db:
await mcp.memory_store(/* … */);  // upsert: true + Promise.all is fine

For >10k entries, prefer bin/build-embeddings.mjs / bin/index-all.mjs — they stream in batches with a progress bar and skip unchanged chunks via a hash file.

Caching

MofloDbAdapter has a built-in LRU cache (default 10k entries, 5-min TTL):

import { MofloDbAdapter } from 'moflo/dist/src/cli/memory/index.js';

const store = new MofloDbAdapter({
  cacheEnabled: true,
  cacheSize: 50_000,      // scale with working set, not total corpus
  cacheTtl: 10 * 60_000,  // 10 minutes
});

Cache hits on exact keys bypass HNSW entirely. If your workload is read-heavy and hits a narrow keyspace, this is the cheapest win.

Measuring

npx vitest bench src/cli/memory/benchmarks/vector-search.bench.ts

The bench prints linear vs HNSW times for 1k and 10k vectors. Run it before and after any parameter change — "it felt faster" is not a benchmark.

For production memory stats:

const stats = await mcp.memory_stats({});
// { entryCount, indexSize, cacheHitRate, avgSearchMs, … }

Common Bottlenecks

SymptomLikely causeFix
Cold-start of 5s on first searchHNSW loading from diskShare a single instance via beforeAll in tests; keep the adapter resident in long-running processes
Search latency climbs linearly with limitOver-fetching and re-ranking on the hot pathLower limit; raise threshold to prune
Inserts slow past ~100k entriesmaxElements too close to entry count → reallocationSet maxElements to 2× expected corpus
High RSS on a small corpusVector dimension mismatch with indexConfirm dimensions matches embedder output (OpenAI = 1536, local models vary)

Anti-Patterns

  • Don't rebuild the index on every test. Use a module-level singleton + beforeAll. HNSW cold-boot is ~5s.
  • Don't raise ef globally. Raise it on the specific queries that need recall. Default is fine for 90% of calls.
  • Don't quantize a small corpus. Below ~500k vectors the RAM saving doesn't justify the recall cost.
  • Don't measure in dev mode. The memory stack behaves differently under NODE_ENV=production; benches should match the target.

See Also

  • memory-patterns skill — API usage and namespace design
  • vector-search skill — RAG-specific patterns on top of the optimized index
  • src/cli/memory/benchmarks/ — runnable benches for every knob above

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

commune

無料

Turn a vague idea into a concrete, actionable spec through a short Socratic dialogue, then hand the result off to an existing moflo surface — a /flo ticket, a spell, or memory. Use BEFORE you have a defined unit of work, when the goal is still fuzzy.

日本語の概要は準備中です。原文の説明を表示しています。

eric-cielo/moflo182026年10月1日 更新

Scaffold new spell step commands and connectors. Use when building new step commands for spells or extending the spell engine with new capabilities. Connectors are for new I/O transport types OR platforms requiring complex multi-step interaction (e.g., browser-based automation).

日本語の概要は準備中です。原文の説明を表示しています。

eric-cielo/moflo182026年10月1日 更新

distill

無料

Alias for /flo-simplify — see that skill's description.

日本語の概要は準備中です。原文の説明を表示しています。

eric-cielo/moflo182026年10月1日 更新

divine

無料

Structured multi-hop web research with explicit confidence gating — plan the inquiry, search (WebSearch/WebFetch), score your own confidence, and keep digging until the answer is well-supported or a hop cap is hit, then emit a cited synthesis. Learns across sessions by storing each research case to memory and reusing prior strategies. Use when a question needs more than one search — comparisons, current-best-practice questions, anything where a single lookup leaves you unsure.

日本語の概要は準備中です。原文の説明を表示しています。

eric-cielo/moflo182026年10月1日 更新

eldar

無料

Consult the Eldar — audit a project's moflo + Claude Code setup for portable, high-leverage gaps and guide remediation. Default mode is read-only audit with severity-ranked findings; --fix presents an interactive triage menu and walks the user through each chosen fix (healer, missing CLAUDE.md, sparse guidance, hook/MCP wiring, empty memory namespaces, stack→guidance gaps). Use when starting in a new project, when Claude feels lost or inefficient, when guidance/CLAUDE.md is sparse, or as a periodic health check.

日本語の概要は準備中です。原文の説明を表示しています。

eric-cielo/moflo182026年10月1日 更新

flfl

無料

Run /fl on a ticket with moflo's three standing considerations loaded first — cross-platform (Rule

日本語の概要は準備中です。原文の説明を表示しています。

eric-cielo/moflo182026年10月1日 更新

eric-cielo のスキルをすべて見る

このスキルの問題を報告する