本文へ移動
cccskills
無料GitHub で公開

metadata-schema

Working with tika-metadata-schema, the build-gated registry of Tika metadata keys — regeneration after Property changes, gate tests, naming conventions, post-rename sweeps. Use when adding or renaming metadata keys or when the schema gate fails.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md6.2 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

<!-- Licensed to the Apache Software Foundation (ASF) under one or more contributor license agreements. See the NOTICE file distributed with this work for additional information regarding copyright ownership. The ASF licenses this file to You under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License. You may obtain a copy of the License at http://www.apache.org/licenses/LICENSE-2.0 Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the specific language governing permissions and limitations under the License. -->

Local override: $TIKA_SKILLS_LOCAL/metadata-schema/LOCAL.md (default ~/.tika-skills), read after this file, wins on conflict.

Metadata Key Registry & Schema Skill

Working with tika-metadata-schema — the committed, build-gated registry of Tika's metadata keys. File structure and the CLOSED/OPEN/TEMPLATE/UNKNOWN classification are in that module's README.md; this covers conventions, regeneration, and the traps.

The three registries

Under tika-metadata-schema/src/main/resources/org/apache/tika/metadata/:

  • metadata-keys.json — closed set: every Property constant + the synthesized tk:digest:* cross-product.
  • metadata-open-namespaces.json — KeyPrefix prefixes for runtime-minted names (html:, message:raw-header:, mdb-prop:).
  • metadata-key-fields.json — TIKA-4797 {class, field, key} table for field-identity migration.

Committed on purpose: they are the reviewable audit trail of the key space (a rename or a dropped key shows up as a diff). Don't switch to build-time-only generation — that loses the review signal.

Regenerate (after adding/changing a Property or KeyPrefix)

tika-metadata-schema/regen.sh

This does the full sequence in one shot: -am install so newly added Property/KeyPrefix classes are on the scan classpath, regenerate all three registries via the forked-exec profile, print a before/after key-count check (catches an incomplete classpath scan), git diff --stat the registries, then run the gate tests. Flags: --skip-install (only safe if nothing outside tika-metadata-schema itself changed since the last install) and --skip-tests for a faster inner loop. Then review the diff and commit the Property change and the regenerated JSON together.

The manual sequence the script replaces, for reference or if you need to run a step in isolation:

# if parser Property classes changed, install them first so the scan sees them:
./mvnw -Pfast -DskipTests -pl tika-metadata-schema -am install -Dmaven.repo.local=$(pwd)/.local_m2_repo
# regenerate all three:
./mvnw -pl tika-metadata-schema -Pregen-metadata-schema process-classes -Dmaven.repo.local=$(pwd)/.local_m2_repo

Trap — never use exec:java. SchemaGenerator scans java.class.path, force-loads Property classes, and swallows load failures. exec:java runs in-process on Maven's classpath → finds zero keys → emits a near-empty registry that exits 0 and passes MetadataNoUnderscoreTest. The -Pregen-metadata-schema profile uses the forking exec goal (<classpath/>) for a real child-JVM classpath. A hand-rolled java -cp also works, but the full repo classpath (~4000 jars) overflows the 128 KB arg limit — use the profile.

Trap — incomplete classpath drops keys silently. Always validate: git diff should show only the intended change (a big key-count drop = classes failed to load), and MetadataSchemaTest regenerates under the full test classpath and asserts a byte-match — the definitive completeness check.

Gate tests (run WITHOUT -Pfast, which skips execution)

./mvnw -pl tika-metadata-schema test -Dmaven.repo.local=$(pwd)/.local_m2_repo
  • MetadataSchemaTest / MetadataFieldTableTest — regenerate in-memory, assert committed files match.
  • MetadataNoUnderscoreTest — no Tika-coined key/prefix may contain _. Scans the JSON, so it reflects the registry, not live code — regenerate before trusting it.
  • MetadataKeyValidatorTest — registry-driven CLOSED/OPEN/TEMPLATE/UNKNOWN classifier.
  • MetadataCoverageTest — fails if a scanned module's keys are neither in scope nor listed out-of-scope.

Failures with stale X-TIKA:/underscore/SHA256 keys usually mean regenerate, not edit code.

Naming conventions (frozen for 4.0, TIKA-4794)

  • All keys are Property constants — no bare String keys (metadata-string-keys.json retired). Sole exception: TikaCoreProperties.EMBEDDED_RESOURCE_TYPE_KEY, an internal building block that constructs the Property name for EMBEDDED_RESOURCE_TYPE — not an independent key.
  • Tika-coined prefix is tk: (X-TIKA: is legacy); kebab-case, no underscores.
  • External-standard names verbatim, including the standard's prefix: dc:, xmp:, cp:, extended-properties:.
  • HTTP has no namespace → Content-Type, Content-Encoding, Location stay bare (no http:).
  • Tika-coined message keys do get a namespace: message:, multipart:.
  • HttpHeaders keys are SIMPLE — Content-Type is not a bag; parsers use set(), not add().
  • Digest keys use the JCA name via DigestDef.getJavaName(): tk:digest:SHA-256, SHA3-256 (not SHA256/SHA3_256). Config input still takes the enum name ("SHA256"); only the output key changes.

After a rename: sweep for stale key literals

The compiler won't catch metadata.get("Message-From"). Grep the whole repo for old strings and prefer replacing them with the constant, so the next rename fails to compile instead of at test time:

grep -rn '"Message-' --include=*.java . | grep -v /target/
grep -rn 'tk:digest:SHA[0-9]' --include=*.java . | grep -v /target/

XML comments can't contain --

A double hyphen in an XML comment makes the POM non-parseable and breaks the module build. Reword.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Ground rules for working in the Tika codebase — git policy, Maven wrapper/repo conventions, building and testing specific modules, code and test conventions, pre-commit checks. Load at session start for any Tika development task.

日本語の概要は準備中です。原文の説明を表示しています。

apache/tika4,0932026年10月10日 更新

Taking a multi-PR feature from "shape unknown" to merged without five review rounds per PR: spike until interfaces stop moving, write the contract, cut PRs along contract seams, one review per PR. Use when starting a feature that touches more than one lifecycle object or public interface, when a PR review keeps changing interfaces, or when splitting a large branch.

日本語の概要は準備中です。原文の説明を表示しています。

apache/tika4,0932026年10月10日 更新

Examine what a file claims about itself and what it actually contains — powered by Apache Tika. True content-based type detection (extensions lie), provenance claims (authors, dates, creating application), revision and tamper signals (PDF incremental updates, tracked changes, hidden slides, zip integrity), hidden and embedded content (attachments, macros), risk indicators (PDF JavaScript actions, encryption), and content digests. Evidence gathering, not verdicts: Tika reports what the file asserts and what parsing observed; it does not attribute authorship or validate signatures. Use for triaging suspicious files, provenance questions, e-discovery-style review, or "is this file what it claims to be."

日本語の概要は準備中です。原文の説明を表示しています。

apache/tika4,0932026年10月10日 更新

Turn almost any file into Markdown plus metadata — PDF, Office, HTML, email, archives, images, audio/video, 1000+ formats — powered by Apache Tika, either via the tika-app CLI (zero setup, one file) or a running tika-server (curl, warm process, many calls). Leads with rmeta (structured, embedded-item-aware output) as the default operation rather than flat concatenated text, since you can't tell whether a file has embedded content from its extension. Covers metadata-only triage, type and language detection, OCR, and output-size discipline for agent context. Use whenever a task involves reading the content of a file whose format you don't want to hand-parse.

日本語の概要は準備中です。原文の説明を表示しています。

apache/tika4,0932026年10月10日 更新

Run Apache Tika as a Docker container when you need guaranteed OCR (scanned PDFs, images) or geospatial raster support with zero local install — `apache/tika:<version>-full` bundles Tesseract, GDAL, ImageMagick, and fonts. Also covers the minimal image, port/volume/memory conventions, the path-identity mount gotcha, and how to confirm OCR actually ran rather than silently returning no text. Powered by Apache Tika. Use when a local `tika-app`/`tika-server` doesn't have Tesseract installed, or you want a disposable, self-contained parsing environment. Companion to the `file-to-markdown` skill, which covers the parsing calls themselves once a server is up.

日本語の概要は準備中です。原文の説明を表示しています。

apache/tika4,0932026年10月10日 更新

oss-fuzz

無料

Run Tika's OSS-Fuzz Jazzer targets locally against a working-tree checkout — build the image, build fuzzers from local source, fuzz a target, run a corpus as a regression pass, reproduce a crash, and add seeds. Use for "fuzz the OneNote parser", "run OneNoteParserFuzzer against these files", "reproduce an OSS-Fuzz crash", "fuzz my branch before merge".

日本語の概要は準備中です。原文の説明を表示しています。

apache/tika4,0932026年10月10日 更新

apache のスキルをすべて見る

このスキルの問題を報告する