本文へ移動
cccskills
無料GitHub で公開

development

Ground rules for working in the Tika codebase — git policy, Maven wrapper/repo conventions, building and testing specific modules, code and test conventions, pre-commit checks. Load at session start for any Tika development task.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md12.2 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

<!-- Licensed to the Apache Software Foundation (ASF) under one or more contributor license agreements. See the NOTICE file distributed with this work for additional information regarding copyright ownership. The ASF licenses this file to You under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License. You may obtain a copy of the License at http://www.apache.org/licenses/LICENSE-2.0 Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the specific language governing permissions and limitations under the License. -->

Local override: $TIKA_SKILLS_LOCAL/development/LOCAL.md (default ~/.tika-skills), read after this file, wins on conflict.

Tika Development Skill

Guidelines and checklist for developing against the Apache Tika codebase.

Question the direction, not just the code

Before optimizing a change — yours or a PR's — ask whether it should exist: does it belong in Tika, is the complexity proportional to the need, would config, an existing mechanism, a plugin, or documentation serve the use case more cheaply? Every merged feature is surface the project maintains for decades. Steelman the use case first and question the vehicle, not the goal; pushback must name a concrete cost or a simpler path — never taste alone, and never no for the sake of no. "The direction is right" is a valid conclusion.

Feature whose shape isn't known yet: spike first, cut PRs after (.skills/devs/feature-workflow/SKILL.md).

Git Policy (default — personally overridable)

Never run git commit or git push — no commits of any kind, including merge commits. If a merge is needed, use git merge --no-commit --no-ff and hand back. Stage files and provide the suggested commit message for the user to run.

Never write to GitHub (PR comments, reviews, issues, labels, merges). Read-only gh is fine.

Precedence: these are conservative defaults for workflow — actions on the contributor's own machine and accounts. A contributor's personal agent configuration (their own skills, CLAUDE.md/AGENTS.md, settings, or a LOCAL.md overlay — see AGENTS.md) may override them. Everything else in this file — code and comment conventions, test discipline, hygiene, pre-commit checks — governs what lands in the repo and is project policy: personal configuration does not override it.

Session Start Checklist

  1. Local Maven repo — Default to an in-repo .local_m2_repo (via -Dmaven.repo.local=$(pwd)/.local_m2_repo) unless the user says otherwise. This isolates builds from the shared ~/.m2/repository and avoids polluting or being affected by other projects.

  2. Maven wrapper — Use ./mvnw; fall back to a system Maven (3.9+) only if the wrapper is absent.

  3. Merge conflicts — Check git status for UU files and resolve before building.

Maven Rules

  • Always include clean in every ./mvnw invocation. Stale classes in target/ cause hard-to-debug failures.

    ./mvnw clean compile -pl <module> ...   # not just: mvnw compile
    ./mvnw clean test -pl <module> ...      # not just: mvnw test
    ./mvnw clean install -pl <module> ...   # not just: mvnw install
    
  • Always use absolute path for local repo:

    -Dmaven.repo.local=$(pwd)/.local_m2_repo
    
  • Fast builds with -Pfast — use the fast profile to skip tests, checkstyle, spotless, and rat in one flag. Prefer this over individual -D skip flags when you want a quick build (e.g., installing for downstream consumers or eval runs):

    ./mvnw clean install -pl <module> -am -Pfast \
      -Dmaven.repo.local=$(pwd)/.local_m2_repo
    

    Run without -Pfast before final commit to catch formatting and style issues. License (rat) checks run only under -Ppedantic (or explicit apache-rat:check), not in default builds.

    -Pfast skips test execution (by design): a green -Pfast build — including -Pfast test — has run zero tests, and stale target/surefire-reports/* will look current. Verify with a plain (non--Pfast) test run.

  • Plugin zips resolve only after package — a reactor clean test fails on modules that depend on pipes plugin zips (tika-server-core, tika-app, ...) unless the zips are already in the local repo. Use clean install (or -Pfast install first).

  • Forked JVM tests — Integration tests in tika-pipes fork new JVMs that load classes from the local Maven repo, not from target/classes. You must ./mvnw clean install -Pfast the changed modules before running integration tests that fork.

Building Specific Modules

# Single module (with dependencies)
./mvnw clean compile -pl <module> -am \
  -Dmaven.repo.local=$(pwd)/.local_m2_repo

# Run a single test class
./mvnw clean test -pl <module> -Dtest=<TestClass> \
  -Dmaven.repo.local=$(pwd)/.local_m2_repo -Dcheckstyle.skip=true

# Install for downstream consumers (tika-app, integration tests)
./mvnw clean install -pl <module> -am -Pfast \
  -Dmaven.repo.local=$(pwd)/.local_m2_repo

Common Module Paths

ModulePath
tika-coretika-core
tika-apptika-app
tika-servertika-server/tika-server-core
tika-evaltika-eval/tika-eval-app
Pipes coretika-pipes/tika-pipes-core
Pipes APItika-pipes/tika-pipes-api
Async CLItika-pipes/tika-async-cli

Code Conventions

  • ASF License 2.0 header on all Java files
  • Spotless formatter runs during build — don't fight it
  • Tests use @TempDir Path tmp for temp directories
  • No emojis in code or comments
  • Comments: every comment must earn its place — one short line by default; multi-line only for a genuinely non-obvious WHY (subtle invariant, workaround, spec quirk). Never restate the code, narrate the next line, justify the change to a reviewer, or describe past states of the code.
  • XML is parsed only through XMLReaderUtils.parseSAX / buildDOM (forbidden-apis rejects the JAXP and Commons Secure XML factories and the raw getSAXParser / getXMLReader / getDocumentBuilder in main code). Never install an EntityResolver; a resolver that returns null or a stream-less InputSource makes the parser fetch the id itself (CVE-2025-66516). A new XML-bearing format gets a fixture in TestXXEInXML; a third-party library that parses XML with its own parser gets its own oracle test (see GeographicInformationXxeTest) before it ships.
  • Input files are hostile: bound anything derived from document content (loop counts, allocations, timeouts); release external processes, temp files, and pool slots on every failure path.
  • No local/machine-specific paths in committed code, tests, docs, or config — never /home/<user>, /Users/<user>, C:\Users\<user>, or a personal ~/data/.... Use a placeholder (<workdir>/, <corpus>), @TempDir, or an in-repo src/test/resources fixture instead. Only legitimate exception: a path that is the data under test (e.g. an expected metadata value extracted from a test document) — leave those untouched.

NEVER: two mistakes that ship as vulnerabilities

Never trust a length that came from user input or from inside a file. Content-Length on a request, a zip or 7z entry's declared size, an OLE/OOXML part size, a PDF embedded file's /Size, an ISO box length, a count line in a mail file: every one of them is an untrusted number. Never allocate, reserve, mark(), skip() or bound a loop by it as if it were true. A lie small truncates content silently; a lie large lets one document allocate the heap or a budget, or spin a reader waiting for bytes that never come. Use a declared length only to decline work (a size the budget cannot cover is not attempted) or as a hint that the actual read corrects; grow from a floor and reserve what was actually read; treat the measured end of the stream as the only ground truth (hasReliableLength()), and overwrite the claim with it. See TIKA-4908 and the 2026-09-19 stream audit for what this cost.

Never call InputStream.skip() on a stream that could be a FileInputStream (or a BufferedInputStream over one) and trust the answer. FileInputStream.skip(n) seeks past EOF and returns n even when no bytes exist, so a reader that skips a payload and then reads the next header believes it is somewhere it is not, and loops or mis-parses forever; BufferedInputStream.skip can return 0 before EOF, which spins any while (remaining > 0) remaining -= skip(remaining) loop. Skip by reading (IOUtils.skip / IOUtils.skipFully) or seek on a channel or an in-memory buffer whose size you own; TikaInputStream.skip, BoundedInputStream.skip and every source's seekTo read for this reason. Do not "optimize" them back. Anything hot enough to need a real seek gets a channel from getSeekableByteChannel(), never a raw skip.

Test Discipline

  • A behavioral change gets a regression test that fails without it. Where impractical (timing, native binaries, external services, kill paths), say so explicitly and name the next-best check.
  • Prove a negative by reverting the fix: an "X does not happen" test that still passes is not a test. Assert on what the consumer is handed, not an ambient side effect (a @TempDir watch misses TemporaryResources not bound to it).
  • Cover error paths and the configuration/mode matrix — a behavior verified in only one parse mode or config shape is a gap (RMETA-only tests miss CONCATENATE-only bugs).
  • Keep tests non-duplicative: don't add a test whose failure another test already guarantees.
  • -Dtika.test.locale=random runs every test JVM under a random default locale (RandomLocaleListener in tika-test-support, wired through tika-parent); run it before touching anything that formats or parses text. The PR jobs leave it off until nightly-locale-build (one full build per priority locale, plus a random draw) has been green several nights running (TIKA-4920). A plain build keeps the JVM's own locale. A failure's stack trace and the fork's stderr name the locale; reproduce with -Dtika.test.locale=<tag>. A test that fails only in some locales is a bug in the code under test (use Locale.ROOT), not a reason to pin Locale.US in the test.
  • Where there's bang for the buck, prefer parameterized tests over copy-pasted cases, randomized inputs over hand-picked ones (log the seed so failures reproduce), and fuzzing for parsers and format/boundary arithmetic. Don't force it on code a couple of fixed cases fully cover.

Metadata Keys & Schema Registry

Adding/renaming a metadata key touches the committed, build-gated registry in tika-metadata-schema — regeneration has real traps. See .skills/devs/metadata-schema/SKILL.md.

Testing an End-to-End Change

When a change affects parsing output (e.g., new parser behavior, encoding fix), run a before/after comparison using tika-eval. See .skills/devs/tika-eval-compare/SKILL.md for the full procedure.

Pre-Commit Checks

# Full compile with checkstyle (catches formatting issues)
./mvnw clean compile -pl <module> -am \
  -Dmaven.repo.local=$(pwd)/.local_m2_repo

# Run module tests
./mvnw clean test -pl <module> \
  -Dmaven.repo.local=$(pwd)/.local_m2_repo

Also before commit:

  • Commit message / PR title references a JIRA ticket (TIKA-XXXX).
  • CHANGES entry for user-visible changes.
  • New dependencies: ASF-compatible license; LICENSE/NOTICE updated.

Scan the staged diff for machine-specific local paths before committing (see Code Conventions). Added lines only; review any hit by hand — a test fixture's expected value is allowed, a real config/doc/code path is not:

git diff --cached -U0 | grep -E '^\+' \
  | grep -nE '/home/[A-Za-z0-9._-]+|/Users/[A-Za-z0-9._-]+|[A-Za-z]:\\+Users|~/data/' \
  && echo "^ local path in staged diff — replace with a placeholder/fixture"

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Taking a multi-PR feature from "shape unknown" to merged without five review rounds per PR: spike until interfaces stop moving, write the contract, cut PRs along contract seams, one review per PR. Use when starting a feature that touches more than one lifecycle object or public interface, when a PR review keeps changing interfaces, or when splitting a large branch.

日本語の概要は準備中です。原文の説明を表示しています。

apache/tika4,0922026年10月10日 更新

Examine what a file claims about itself and what it actually contains — powered by Apache Tika. True content-based type detection (extensions lie), provenance claims (authors, dates, creating application), revision and tamper signals (PDF incremental updates, tracked changes, hidden slides, zip integrity), hidden and embedded content (attachments, macros), risk indicators (PDF JavaScript actions, encryption), and content digests. Evidence gathering, not verdicts: Tika reports what the file asserts and what parsing observed; it does not attribute authorship or validate signatures. Use for triaging suspicious files, provenance questions, e-discovery-style review, or "is this file what it claims to be."

日本語の概要は準備中です。原文の説明を表示しています。

apache/tika4,0922026年10月10日 更新

Turn almost any file into Markdown plus metadata — PDF, Office, HTML, email, archives, images, audio/video, 1000+ formats — powered by Apache Tika, either via the tika-app CLI (zero setup, one file) or a running tika-server (curl, warm process, many calls). Leads with rmeta (structured, embedded-item-aware output) as the default operation rather than flat concatenated text, since you can't tell whether a file has embedded content from its extension. Covers metadata-only triage, type and language detection, OCR, and output-size discipline for agent context. Use whenever a task involves reading the content of a file whose format you don't want to hand-parse.

日本語の概要は準備中です。原文の説明を表示しています。

apache/tika4,0922026年10月10日 更新

Run Apache Tika as a Docker container when you need guaranteed OCR (scanned PDFs, images) or geospatial raster support with zero local install — `apache/tika:<version>-full` bundles Tesseract, GDAL, ImageMagick, and fonts. Also covers the minimal image, port/volume/memory conventions, the path-identity mount gotcha, and how to confirm OCR actually ran rather than silently returning no text. Powered by Apache Tika. Use when a local `tika-app`/`tika-server` doesn't have Tesseract installed, or you want a disposable, self-contained parsing environment. Companion to the `file-to-markdown` skill, which covers the parsing calls themselves once a server is up.

日本語の概要は準備中です。原文の説明を表示しています。

apache/tika4,0922026年10月10日 更新

Working with tika-metadata-schema, the build-gated registry of Tika metadata keys — regeneration after Property changes, gate tests, naming conventions, post-rename sweeps. Use when adding or renaming metadata keys or when the schema gate fails.

日本語の概要は準備中です。原文の説明を表示しています。

apache/tika4,0922026年10月10日 更新

oss-fuzz

無料

Run Tika's OSS-Fuzz Jazzer targets locally against a working-tree checkout — build the image, build fuzzers from local source, fuzz a target, run a corpus as a regression pass, reproduce a crash, and add seeds. Use for "fuzz the OneNote parser", "run OneNoteParserFuzzer against these files", "reproduce an OSS-Fuzz crash", "fuzz my branch before merge".

日本語の概要は準備中です。原文の説明を表示しています。

apache/tika4,0922026年10月10日 更新

apache のスキルをすべて見る

このスキルの問題を報告する