本文へ移動
cccskills
無料GitHub で公開

tika-eval-encoding-regression

Condensed tika-eval pattern for charset-detector regression hunts ("A picks encoding X, B picks Y") using one build and two configs — encoding-pair flip queries, OOV/languageness/FFFD signals, per-file detector attribution.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md9.6 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

<!-- Licensed to the Apache Software Foundation (ASF) under one or more contributor license agreements. See the NOTICE file distributed with this work for additional information regarding copyright ownership. The ASF licenses this file to You under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License. You may obtain a copy of the License at http://www.apache.org/licenses/LICENSE-2.0 Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the specific language governing permissions and limitations under the License. -->

Local override: $TIKA_SKILLS_LOCAL/tika-eval-encoding-regression/LOCAL.md (default ~/.tika-skills), read after this file, wins on conflict.

tika-eval for encoding-detector regression hunts

A condensed pattern for finding SBCS→CJK style charset-detector regressions (or any "A picks encoding X, B picks encoding Y" question) without building two tika-app distributions.

Two configs, one build

Encoding-detector experiments don't need a "before" and "after" tika-app — the chain composition is per-config. Run the SAME tika-app twice against two configs, treat the outputs as -a and -b. Much faster than tika-eval-compare's two-build flow.

# build once
./mvnw clean install -pl tika-app -am -Pfast -DskipTests \
  -Dmaven.repo.local=$(pwd)/.local_m2_repo
unzip -q tika-app/target/tika-app-*.zip -d /tmp/tika-app-current

# two configs (any combination of detectors)
java -jar /tmp/tika-app-current/tika-app-*.jar \
  --config=tika-config-3x-default.json \
  -i <corpus> -o <workdir>/extracts/A -n 6
java -jar /tmp/tika-app-current/tika-app-*.jar \
  --config=tika-config-junkfilter-combiner.json \
  -i <corpus> -o <workdir>/extracts/B -n 6

# normal Compare
java -jar /tmp/tika-eval-current/tika-eval-app-*.jar Compare \
  -a <workdir>/extracts/A -b <workdir>/extracts/B -d <workdir>/extracts/A-vs-B -r -rd <workdir>/extracts/A-vs-B-reports

Canonical 3.x-default encoding chain config

{
  "encoding-detectors": [
    {"html-encoding-detector": {}},
    {"universal-encoding-detector": {}},
    {"icu4j-encoding-detector": {}}
  ]
}

Canonical 4.x junkfilter chain config

{
  "encoding-detectors": [
    {"bom-detector": {}},
    {"html-encoding-detector": {}},
    {"mojibuster-encoding-detector": {}},
    {"junk-filter-encoding-detector": {}}
  ]
}

Per-detector isolation configs

Each detector wired alone lives in <workdir>/configs/: tika-config-bom.json, tika-config-html.json, tika-config-htmlstandard.json, tika-config-universal.json, tika-config-icu4j.json, tika-config-mojibuster.json, tika-config-junkfilter-chain.json. Use these for chain-attribution work (which detector did the detection).

Encoding-pair flip query

MIMES.MIME_STRING for text-y mimes is text/html; charset=X form. Extract the charset with a regex split, group by (enc_a, enc_b), filter pairs. A=before/-a, B=after/-b; join on pa.ID = pb.ID (paired by id).

SELECT
  REGEXP_REPLACE(ma.MIME_STRING, '^.*charset=', '') AS enc_a,
  REGEXP_REPLACE(mb.MIME_STRING, '^.*charset=', '') AS enc_b,
  COUNT(*) n,
  SUM(cb.NUM_COMMON_TOKENS - ca.NUM_COMMON_TOKENS) AS delta_common
FROM PROFILES_A pa
JOIN PROFILES_B pb ON pa.ID = pb.ID
JOIN MIMES ma ON pa.MIME_ID = ma.MIME_ID
JOIN MIMES mb ON pb.MIME_ID = mb.MIME_ID
JOIN CONTENTS_A ca ON ca.ID = pa.ID
JOIN CONTENTS_B cb ON cb.ID = pb.ID
WHERE ma.MIME_STRING LIKE '%charset=%' AND mb.MIME_STRING LIKE '%charset=%'
  AND REGEXP_REPLACE(ma.MIME_STRING, '^.*charset=', '') <>
      REGEXP_REPLACE(mb.MIME_STRING, '^.*charset=', '')
GROUP BY enc_a, enc_b
ORDER BY n DESC, delta_common ASC LIMIT 50;

Add an IN (...) filter on either side to constrain to a family (e.g. SBCS-Western → CJK):

  AND REGEXP_REPLACE(ma.MIME_STRING,'^.*charset=','')
      IN ('windows-1252','ISO-8859-1','ISO-8859-15','ISO-8859-2','ISO-8859-3',
          'windows-1250','windows-1254','windows-1257','ISO-8859-13',
          'windows-1258','x-MacRoman','IBM850','IBM852')
  AND REGEXP_REPLACE(mb.MIME_STRING,'^.*charset=','')
      IN ('GB18030','GBK','GB2312','Big5','Big5-HKSCS','Shift_JIS','EUC-JP',
          'EUC-KR','x-EUC-TW','x-windows-874','x-windows-949',
          'ISO-2022-JP','ISO-2022-KR','ISO-2022-CN')

Per-file drilldown

Join CONTAINERS to get the source path; pull LANG_ID_1 from both sides to see whether language detection agrees the content is Western while the charset has flipped to CJK (the regression's defining shape):

SELECT ct.FILE_PATH,
       REGEXP_REPLACE(ma.MIME_STRING,'^.*charset=','') AS enc_a,
       REGEXP_REPLACE(mb.MIME_STRING,'^.*charset=','') AS enc_b,
       ca.NUM_COMMON_TOKENS AS ca_tok, cb.NUM_COMMON_TOKENS AS cb_tok,
       cb.NUM_COMMON_TOKENS - ca.NUM_COMMON_TOKENS AS delta,
       ca.LANG_ID_1 AS lang_a, cb.LANG_ID_1 AS lang_b
FROM PROFILES_A pa JOIN PROFILES_B pb ON pa.ID = pb.ID
JOIN MIMES ma ON pa.MIME_ID = ma.MIME_ID JOIN MIMES mb ON pb.MIME_ID = mb.MIME_ID
JOIN CONTENTS_A ca ON ca.ID = pa.ID JOIN CONTENTS_B cb ON cb.ID = pb.ID
JOIN CONTAINERS ct ON ct.CONTAINER_ID = pa.CONTAINER_ID
WHERE <enc_a/enc_b filter as above>
ORDER BY delta ASC LIMIT 15;

Reading the signals — OOV, languageness, and FFFD together

No single signal is authoritative. Use oov as a secondary signal alongside languageness (the junk-model coherence z-score) and the U+FFFD rate — each is right where the others are blind, so cross-check rather than ranking on any one. (Established 2026-06-03: a 40-file OOV-"worse" set was mostly metric artifacts once languageness/FFFD were brought in — only ~6 were real. But OOV is also the correct signal where languageness is blind, so neither dominates.)

  • OOV can mislead when langid shifts — a CJK/UTF-8 recovery in B is scored against a different vocab → higher OOV though B is right — or when a wrong decode fragments words into more short common tokens (→ higher count for the WORSE decode). A common-token delta is a signal, not proof.
  • languageness can mislead on SBCS↔SBCS cross-script mojibake — Greek decoded as KOI8-R is "coherent" Cyrillic, so languageness stays flat while oov correctly flags it. Conversely languageness catches OOV's CJK/script-recovery blind spot. Each covers the other's blind spot.
  • FFFD rate flags decode failures (illegal bytes): num_replacement / num_non_ascii (un-diluted; / content_length dilutes to ~0 on ASCII-heavy docs). Tika strips C0 controls at extraction, so legal-but-wrong (C1) mojibake does not surface here — that signal belongs in the detector chain, not the eval.
  • In practice: when the signals agree, high confidence; when they disagree (OOV-worse but languageness-better, or vice versa), that file needs a look — the disagreement points you at WHICH files to inspect, it does not by itself declare OOV or languageness "wrong." Split OOV-worse by languageness direction (query in tika-eval-regression.adoc).

Isolate a change against the PRIOR run, not just 3.x

To see what one chain change actually did, Compare the new run against the previous 4.x run (B-new vs B-prior), not only vs 3.x. The diff should be surgical — e.g. the within-Latin letter gate moved exactly 6 files (IBM850 / x-MacRoman → windows-1252) vs the prior run and nothing else. A bigger-than-expected diff means the change fired more broadly than intended.

Per-file detector attribution (tk:encoding-detection-trace)

Every JSON extract from a chain with multiple detectors carries tk:encoding-detection-trace in metadata. It's a per-detector emission log with the META detector's arbitration tag at the end:

MojibusterEncodingDetector->Shift_JIS[STATISTICAL](1.00) [junk-filter-selected]

When investigating "why did B pick X for this file?", read this trace first — it tells you which base detector(s) emitted candidates and which one the meta detector chose. If the trace shows ONLY Mojibuster firing with a CJK pick, the bug is in Mojibuster's emission (pool too narrow), not in JunkFilter's arbitration.

tk:encoding-detector is the simple-name credit string; tk:detected-encoding is the final answer (also in Content-Encoding).

Reproducing a single-file detection without a full chain

./mvnw -q -pl tika-ml/tika-ml-junkdetect -Dmaven.repo.local=$(pwd)/.local_m2_repo \
  -Dexec.classpathScope=test \
  -Dexec.mainClass=org.apache.tika.ml.junkdetect.TraceJunkFilter \
  -Dexec.args="--file <path> --auto-candidates --content-cleaner --head-bytes 524288 --sample 120" \
  exec:java

Key flags:

  • --auto-candidates — use Mojibuster's per-file pool as the candidate set
  • --content-cleaner — decode each candidate then run text through HtmlContentCleaner to match the live chain
  • --head-bytes 524288 — read up to 512 KB raw to match AdaptiveProbe.DEFAULT_RAW_CAP. The default READ_LIMIT of 16 KB will give a different probe than the live chain on long markup-heavy pages and lead you to disagree with the live chain's pick. Always pass this when reconciling a TraceJunkFilter run with a live extract.

Without --head-bytes, you are looking at a different probe than the chain saw — this is the most common source of "trace says X, chain says Y" confusion.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Ground rules for working in the Tika codebase — git policy, Maven wrapper/repo conventions, building and testing specific modules, code and test conventions, pre-commit checks. Load at session start for any Tika development task.

日本語の概要は準備中です。原文の説明を表示しています。

apache/tika4,0932026年10月10日 更新

Taking a multi-PR feature from "shape unknown" to merged without five review rounds per PR: spike until interfaces stop moving, write the contract, cut PRs along contract seams, one review per PR. Use when starting a feature that touches more than one lifecycle object or public interface, when a PR review keeps changing interfaces, or when splitting a large branch.

日本語の概要は準備中です。原文の説明を表示しています。

apache/tika4,0932026年10月10日 更新

Examine what a file claims about itself and what it actually contains — powered by Apache Tika. True content-based type detection (extensions lie), provenance claims (authors, dates, creating application), revision and tamper signals (PDF incremental updates, tracked changes, hidden slides, zip integrity), hidden and embedded content (attachments, macros), risk indicators (PDF JavaScript actions, encryption), and content digests. Evidence gathering, not verdicts: Tika reports what the file asserts and what parsing observed; it does not attribute authorship or validate signatures. Use for triaging suspicious files, provenance questions, e-discovery-style review, or "is this file what it claims to be."

日本語の概要は準備中です。原文の説明を表示しています。

apache/tika4,0932026年10月10日 更新

Turn almost any file into Markdown plus metadata — PDF, Office, HTML, email, archives, images, audio/video, 1000+ formats — powered by Apache Tika, either via the tika-app CLI (zero setup, one file) or a running tika-server (curl, warm process, many calls). Leads with rmeta (structured, embedded-item-aware output) as the default operation rather than flat concatenated text, since you can't tell whether a file has embedded content from its extension. Covers metadata-only triage, type and language detection, OCR, and output-size discipline for agent context. Use whenever a task involves reading the content of a file whose format you don't want to hand-parse.

日本語の概要は準備中です。原文の説明を表示しています。

apache/tika4,0932026年10月10日 更新

Run Apache Tika as a Docker container when you need guaranteed OCR (scanned PDFs, images) or geospatial raster support with zero local install — `apache/tika:<version>-full` bundles Tesseract, GDAL, ImageMagick, and fonts. Also covers the minimal image, port/volume/memory conventions, the path-identity mount gotcha, and how to confirm OCR actually ran rather than silently returning no text. Powered by Apache Tika. Use when a local `tika-app`/`tika-server` doesn't have Tesseract installed, or you want a disposable, self-contained parsing environment. Companion to the `file-to-markdown` skill, which covers the parsing calls themselves once a server is up.

日本語の概要は準備中です。原文の説明を表示しています。

apache/tika4,0932026年10月10日 更新

Working with tika-metadata-schema, the build-gated registry of Tika metadata keys — regeneration after Property changes, gate tests, naming conventions, post-rename sweeps. Use when adding or renaming metadata keys or when the schema gate fails.

日本語の概要は準備中です。原文の説明を表示しています。

apache/tika4,0932026年10月10日 更新

apache のスキルをすべて見る

このスキルの問題を報告する