MIME Detection & Routing
Detection Flow
Policy -> content and/or extension evidence -> validate_mime_type -> registry.get(mime) -> extractor
Key Functions
| Function | Location | Behaviour |
|---|---|---|
detect_mime_type(path, check_exists: bool) |
core/mime.rs |
Path-based only — never reads bytes. Lowercased extension → EXT_TO_MIME, then tree-sitter extension detection (feature tree-sitter), then mime_guess::from_path. check_exists gates a file-existence check, not content inspection. |
detect_mime_type_from_bytes(bytes) |
core/mime.rs |
Magic-number detection via the infer crate. The only content-sniffing entry point. |
validate_mime_type(mime) |
core/mime.rs |
Parses the media type, matches its case-insensitive essence against SUPPORTED_MIME_TYPES, and returns the registered MIME spelling. Parameters such as charset do not affect extractor routing. It does not consult the extractor registry. |
The FORMATS registry is the single source of truth
FORMATS: &[FormatEntry { extensions, mime_type, aliases }] in core/mime.rs.
EXT_TO_MIME and SUPPORTED_MIME_TYPES are LazyLocks derived from it by iteration —
there is no m.insert call site to add to, and hand-editing either is impossible.
The full registry publishes 106 formats, 140 unique extensions, and 53 aliases, verified by
scripts/sync_supported_counts.py verify. The published count constants describe that static
registry; runtime availability is its intersection with registered extractors. Extension lookup is
case-insensitive (the extension is lowercased before the map hit).
Registry Selection
let registry = get_document_extractor_registry(); // plugins/registry/mod.rs
let guard = registry.read()?;
let extractor: Arc<dyn DocumentExtractor> = guard.get(mime_type)?; // Result, not Option
DocumentExtractorRegistry::get (plugins/registry/extractor.rs) returns the
highest-priority() extractor for the MIME type, and returns Err — not None — when none
matches.
Wildcard Support
An extractor may register a family: "image/*" matches image/png, image/jpeg, and so on
(prefix match on a registered type ending in /*).
Adding a New Format
- Add one
FormatEntrytoFORMATSincrates/xberg/src/core/mime.rs.EXT_TO_MIMEandSUPPORTED_MIME_TYPESupdate automatically. - Run
scripts/sync_supported_counts.py syncto update published count claims, then run itsverifycommand. - Implement
InternalDocumentExtractor(notDocumentExtractor— seeplugin-architecture-patterns) withsupported_mime_types()returning the MIME. - Register in
crates/xberg/src/extractors/mod.rs::register_default_extractors().
Critical Rules
- Call
validate_mime_type()before extraction — but do not treat it as proof an extractor exists. - Extension lookup is case-insensitive.
detect_mime_typeinspects no content. Extraction defaults toPreferContent, which performs bounded content inspection and falls back to a supported extension. UseContentOnlywhen the filename must be ignored; useTrustExtensiononly for trusted sources.- A specific explicit MIME type is authoritative.
application/octet-streamis the exception: it is a generic placeholder and triggers policy-based detection. - Never edit
EXT_TO_MIMEorSUPPORTED_MIME_TYPES— editFORMATS.