Skip to content

Raw extraction formats

Reference for the extraction stage — how each raw source under raw/<key>/ is read and mapped onto the unified Document schema in extracted/. Use this when auditing an extractor or adding a new one.

For the commands and the rationale behind the text/annotations split, see Extract. This page documents the formats, not the CLI.

The stage in one pass

scripts/data/extract.py is a thin wrapper around slm4ie.data.processing.extract_datasets. Every dataset key in configs/data/extract.yaml declares two things:

  • extractor — the registry name of the reader (the eight documented below).
  • domain — a free-text provenance tag (web, wiki, parliamentary, academic, legal, medical, scientific, forum, blog, news, student, mixed) that drives source-weighted sampling downstream.

Per dataset the orchestrator then:

  1. Decompresses any archive in raw/<key>/ — suffixes .gz, .xz, .zip, .tgz, .tar.gz, .tar.zst, .tar.zstd.
  2. Dispatches to the registered extractor, which yields Document objects.
  3. Writes two artifacts, never merged:
    • extracted/<key>.jsonltext, source, domain, doc_id, uid, and metadata.
    • extracted/<key>.annotations.jsonl.gz — gzipped parallel arrays (forms, lemmas, upos, feats, sentences).

The annotations sidecar is written only if at least one document in the dataset carries real annotations. In a mixed dataset, unannotated documents get a stub line (doc_id + uid only) so the two files stay line-aligned; readers detect a stub by the absence of the parallel-array fields. Stubs seen before the first real annotation are buffered and flushed once one appears — so a fully unannotated dataset produces no sidecar at all.

Both files are written to .partial and promoted with os.replace, so a crashed run never leaves a half-written output in place. Outputs are skipped unless --force is passed.

The unified schema

Defined in slm4ie/data/schema.py:

Document(text, source, domain, doc_id, metadata, annotations)
Annotations(tokens=[Token(form, lemma, upos, feats)], sentences=[[start, end], ...])

uid is a derived property — "{source}:{doc_id}" — so documents from different corpora can never collide on a reused internal id. Sentence spans in Annotations.sentences are inclusive [start, end] token-index pairs over the flat token list: a two-sentence document with 5 + 7 tokens has sentences == [[0, 4], [5, 11]].

Parallelism

Datasets are always processed sequentially; --max-workers sets shard workers used within one dataset. Sharding only engages for extractors that subclass FileBasedExtractor and have at least 8 input files and more than one worker. Shards are parsed into temp files and concatenated in order, so the merged output is byte-identical to the serial writer's.

Extractor Datasets Annotations? Shardable
jsonl classla_web_sl, legal_mc4, slovenian_news Only with paragraphs No
json povejmo_vemo_med No No
text cc100 No Yes
conllu classlawiki_sl, oss, kzb, solar, suk, ssj500k Yes Yes
tei parlamint_si, kas, siparl, janes_forum, janes_blog, janes_news, gigafida Yes, on the annotated paths Yes
macocu macocu_sl No Yes
coleslaw coleslaw No No
huggingface finepdf, fineweb2, culturax, c4, hplt No No

All extractors discover their input files recursively (rglob) in sorted order. Registry and base classes live in slm4ie/data/extractors/__init__.py.

jsonl

slm4ie/data/extractors/jsonl.py — line-delimited JSON (*.jsonl), one object per line. Blank and malformed lines are skipped with a warning.

{"doc_id": "d1", "text": "Dober dan.", "url": "https://example.com",
 "paragraphs": [{"sentences": [{"tokens": [
   {"form": "Dober", "lemma": "dober", "upos": "ADJ", "feats": "Case=Nom"},
   {"form": "dan", "lemma": "dan", "upos": "NOUN", "feats": "Case=Nom"},
   {"form": ".", "lemma": ".", "upos": "PUNCT", "feats": null}]}]}]}
Raw Document
text (configurable) text — records with empty/missing text are skipped
doc_id (configurable) doc_id
paragraphs annotations, flattened across paragraphs → sentences → tokens
all other fields metadata

The field names are configurable through the metadata: block, so a feed that names its fields differently needs no bespoke extractor. slovenian_news uses this:

slovenian_news:
  extractor: jsonl
  domain: news
  metadata:
    text_field: body            # default: "text"
    id_field: uri               # default: "doc_id"
    metadata_fields: [url, title, dateTime, source]

When metadata_fields is omitted every record field is kept as metadata except the text field, the id field, and the structural fields paragraphs and conll. When given, only the listed keys are kept.

Annotations are produced only when a paragraphs field is present and yields at least one token — classla_web_sl carries them, legal_mc4 and slovenian_news do not.

json

slm4ie/data/extractors/json.py*.json holding a top-level array. A single top-level object is accepted and treated as a one-record array; anything else is skipped with a warning.

[
  {"doc_id": "vemo.1", "text": "Bolnik je prišel z bolečinami.",
   "specialty": "interna", "year": 2023},
  {"doc_id": "vemo.2", "text": "Drugi opis primera."}
]
Raw Document
text (configurable) text — empty/missing skipped
doc_id (configurable) doc_id
every other non-null field metadata

Text-only; no annotations. The field names are configurable through the metadata: block with the same knobs as jsonltext_field (default text), id_field (default doc_id), and metadata_fields (whitelist of record fields to keep; when omitted, every other field is kept). Fields with a null value are always dropped from metadata.

text

slm4ie/data/extractors/text.py*.txt streamed line-by-line so multi-GB inputs never need to fit in memory. A blank line is a document boundary (the CC100 convention).

Prvi dokument, prva vrstica.
Prvi dokument, druga vrstica.

Drugi dokument, ena sama vrstica.

Tretji dokument.

text is the block's non-empty lines joined with newlines and stripped.

Raw Document
block's non-empty lines text, joined with newlines and stripped
file path + block position doc_id<rel>:<block index>

rel is the file's path relative to raw/<key>/ with its suffix dropped and separators normalized to /; the block index is 0-based and restarts at zero for each file, zero-padded to six digits. So raw/cc100/sl.txt yields sl:000000, sl:000001, … and uid becomes cc100:sl:000000. Blocks that strip to nothing are skipped without consuming an index, so ids stay contiguous. See Deterministic ids for why the id is derived this way.

No metadata and no annotations.

conllu

slm4ie/data/extractors/conllu.py*.conllu and *.conll. Ten tab-separated columns, # comment lines, blank line = sentence boundary, _ = missing value.

# newdoc id = doc1
# sent_id = doc1.s1
# text = Predsednik je odprl sejo.
1   Predsednik  predsednik  NOUN    Ncmsn   Case=Nom    3   nsubj   _   NER=O
2   je  biti    AUX Va-r3s-n    Tense=Pres  3   aux _   NER=O
3   odprl   odpreti VERB    Vmep-sm VerbForm=Part   0   root    _   NER=O
4   sejo    seja    NOUN    Ncfsa   Case=Acc    3   obj _   NER=O
5   .   .   PUNCT   Z   _   3   punct   _   SpaceAfter=No

Of the ten columns, only four become Token fields:

Column Token
2 FORM form
3 LEMMA lemma
4 UPOS upos
6 FEATS feats
10 MISC SpaceAfter=Nospace_after: false; nothing else is read

ID, XPOS, HEAD, DEPREL, and DEPS are not carried into the schema. Rows whose ID marks a multiword token (1-2) or an empty node (1.1) are skipped.

Document boundaries are detected in priority order:

  1. # newdoc id = ... markers — the standard CoNLL-U signal.
  2. A change in the leading component of a hierarchical # sent_id (solar1.1.1solar2.1.1 opens a document with doc_id = "solar2"). Sources like Solar pack thousands of essays into one file without ever emitting # newdoc id. This heuristic is disabled for the rest of the file as soon as any # newdoc id is seen.
  3. One document per file, doc_id = filename stem.

Text prefers each sentence's # text = comment; when absent it is reconstructed from the tokens, honouring SpaceAfter=No. Sentence strings are then joined with newlines. oss, kzb, and kas pull extra per-document fields from a metadata sidecar.

Spacing is persisted, not just used for reconstruction: each token's Token.space_after is False when its MISC field contains SpaceAfter=No, and the annotations sidecar carries a space_after bool array parallel to forms. Both the reconstructed text and the array derive from the same MISC signal, so they cannot diverge. Sidecars written before this field existed lack the array — consumers treat a missing array as all-True. Re-extract with --force to gain it.

tei

slm4ie/data/extractors/tei.py*.xml in the TEI namespace http://www.tei-c.org/ns/1.0. Three structural paths are auto-detected from the tree itself:

Structure Detected by Unit
Annotated with utterances has <w>, has <u> one document per <u>
Annotated without utterances has <w>, no <u> one document per file
Plain no <w> one document per <p>
<TEI xmlns="http://www.tei-c.org/ns/1.0">
  <text><body>
    <u xml:id="u1" who="#chair">
      <s xml:id="u1.s1">
        <w lemma="dober" msd="UPosTag=ADJ|Case=Nom">Dober</w>
        <w lemma="dan" msd="UPosTag=NOUN|Case=Nom">dan</w>
        <pc msd="UPosTag=PUNCT">.</pc>
      </s>
    </u>
  </body></text>
</TEI>
Raw Document
<w> / <pc> element text Token.form
lemma attribute Token.lemma
msd or ana attribute Token.upos + Token.feats
join attribute Token.space_after (see below)
xml:id on <u> (or <p>) doc_id; falls back to the filename stem on the per-file path
who / ana on <u> metadata, layered over any sidecar fields

Tokens are collected from <w> and <pc> children of each <s>, including those wrapped in a <name> element. Morphology is read from msd first (UPosTag=X|Key=Val → the UPosTag part becomes upos, the rest is rejoined with | as feats); when absent, ana is parsed for an mte: MULTEXT-East v6 compact code, which is mapped to UPOS by its category character (refined by the second character for NOUN/PROPN, VERB/AUX, CCONJ/SCONJ) and preserved verbatim as feats="MTE=<code>".

On the plain path, <p> scanning is scoped to <body> so that <p> elements in the <teiHeader> (copyright notices, catalog codes) are never emitted as documents; text comes from .itertext().

Spacing comes from the TEI/CLARIN.SI join attribute on <w>/<pc>, which marks the absence of whitespace: right = no space to the token's right, left = no space to its left, both = neither side. KAS and ParlaMint-SI place join="right" on the token preceding punctuation (<w join="right">uprava</w><pc>.</pc>). Between consecutive tokens A, B the space is dropped when A is right-joined or B is left-joined; the result is stored as Token.space_after on A and persisted in the sidecar's space_after array, and reconstructed text honours it — javna uprava., not javna uprava .. This matches the conllu extractor's SpaceAfter=No handling, so reconstructed-text fidelity is uniform across the annotated routes. Sentence strings are still joined with newlines regardless of the final token's flag. Sidecars extracted before this field existed lack the array (consumers treat missing as all-True) and carry space-joined text — re-extract with --force to refresh both.

Files larger than 64 MB are parsed with iterparse(huge_tree=True) instead of a full DOM — GigaFida ships segments whose DOM would be 10–15× the file size and OOM the extractor under sharding. The streaming path only handles the per-file annotated structure; utterance- and plain-structured files fall back to the full-DOM path.

macocu

slm4ie/data/extractors/macocu.py — the MaCoCu monolingual DTD (*.xml), streamed with lxml iterparse filtered on <doc> so large corpora never build a full tree.

<corpus id="MaCoCu-sl-2.0">
  <doc id="macocu.sl.1" title="Page One" url="https://example.com/1"
       crawl_date="2022-07-01" lm_score="0.95">
    <p id="macocu.sl.1.1" lang="sl">Dober dan.</p>
    <p id="macocu.sl.1.2" lang="sl">Kako ste?</p>
  </doc>
</corpus>
Raw Document
<p> contents text, joined with newlines; empty <doc> skipped
id attribute on <doc> doc_id
title, crawl_date, lang_distr, url, domain, file_type, lm_score on <doc> metadata (only those present and non-empty)

Text-only; no annotations. <doc> attributes outside that whitelist are dropped.

coleslaw

slm4ie/data/extractors/coleslaw.py*.jsonl across four legal subcorpora that ship different schemas, detected per record by field presence.

{"id": 1, "text": "Zakon o nečem.", "title": "Zakon"}
{"id": "Up-1", "fullText": "Sklep ustavnega sodišča."}
{"id": "c1", "jedro": "Bistvo.", "izrek": "Razveljavi se.", "obrazlozitev": "Obrazložitev sledi."}
{"id": "750", "skodni_dogodek": "Prometna nesreča.", "poskodba": "Zvin vratu."}

text is the first non-empty of, in order:

  1. text — PISRS, UradniList.
  2. fullText — USRS (Constitutional Court).
  3. jedro, izrek, obrazlozitev joined with blank lines — SodnaPraksa sp_courts. The order mirrors the structure of Slovenian court decisions: essence → operative part → reasoning.
  4. Six sp_claims prose fields (skodni_dogodek, poskodba, telesne_bolecine, strah, zmanjsanje_zivljenjske_aktivnosti, dodatne_informacije) joined in reading order — personal-injury summaries have no unified text body.

The precedence assumes the four subcorpora expose mutually exclusive text fields. If a record presents more than one of these sources, the earliest still wins, but the extractor logs a warning naming the record's doc_id, subcorpus, all present sources, and the chosen one — a shadowed field is never silent.

doc_id is doc_id if present, else id coerced to a string. metadata is {"subcorpus": <parent directory name>}PISRS, UradniList, SodnaPraksa, USRS — plus every remaining non-null field that was not consumed for text. Note that id is not reserved, so it survives into metadata even when it supplied the doc_id. Text-only; no annotations.

huggingface

slm4ie/data/extractors/huggingface.py — each immediate subdirectory of raw/<key>/ is an Arrow dataset written by save_to_disk() and loaded with load_from_disk(). Both Dataset and DatasetDict are handled; for a DatasetDict every split is iterated. A config dir that fails to load is logged and skipped.

raw/c4/
  sl/                       # config 1 — Slovene
    dataset_info.json
    state.json
    data-00000-of-00001.arrow
  hr/                       # config 2 — Croatian
    ...

A row, e.g. for AllenAI C4:

{"text": "Dober dan, kako ste?",
 "timestamp": datetime(2019, 4, 25, 12, 34, 56),
 "url": "https://example.com/page"}
Raw Document
text column text — empty/missing rows skipped
first of id, _id, docid, doc_id, uid, url doc_id, coerced to a string
every other non-null column metadata

datetime and date values (matched by Python type, not column name) are converted to ISO-8601 strings so the row stays JSON-serializable. No annotations.

The key column is probed per row in the priority order above (the NATURAL_ID_KEYS constant), taking the first present, non-empty value. It is also kept in metadata — using a column as the id never removes it. Rows with no such column fall back to a positional <config>:<split>:<row index> — the config subdirectory name, the split name (omitted for a bare Dataset), and the 0-based row index within the split, zero-padded to eight digits. So a keyless raw/c4/sl/ yields sl:00000000, and a DatasetDict yields sl:train:00000000.

The row index counts every row, including those skipped for empty text, so an id stays pinned to its offset in the source split. Note that natural-key uniqueness is the source's to guarantee: a corpus with duplicate url values would produce duplicate uids.

External-TSV metadata

conllu and tei can merge per-document fields from a flat TSV shipped alongside the text files, implemented in slm4ie/data/metadata_sidecar.py. Used by oss, kzb, and kas.

The lookup key is derived from the filename stem, optionally narrowed by a key_pattern regex whose first capture group becomes the actual key. The matched row is projected through fields (which also renames columns) and splits (which explodes separator-joined columns into JSON arrays). Values of "" or "-" are dropped as NA. The resulting dict is applied to every document produced from that file, since rows are keyed by filename rather than by document.

Given oss-10000.conllu and this config:

oss:
  extractor: conllu
  domain: scientific
  metadata:
    path: OSS.CoNLL-U/OSS-metadata.tsv
    key_column: id
    key_from: filename_stem
    key_pattern: '^oss-(\d+)$'    # stem "oss-10000" → key "10000"
    fields:
      cerif: cerif                # tsv column → metadata field
      udc: udc
      type: doctype               # renamed on the way in
    splits:
      cerif: "|"                  # explode into a list

and this TSV:

id  cerif   udc type
10000   P000|T270   502(043)    Diplomsko delo

every document from that file gets:

{"cerif": ["P000", "T270"], "udc": "502(043)", "doctype": "Diplomsko delo"}

key_from currently accepts only filename_stem. On the TEI utterance path, who / ana take precedence over sidecar fields on a key collision, being the more specific source.

Notes

Cross-cutting behaviour worth knowing before extending the stage.

Only three routes produce annotations

conllu, tei (on its two annotated paths), and jsonl (when a record carries paragraphs). The other five — json, text, macocu, coleslaw, huggingface — are text-only, and their datasets get no .annotations.jsonl.gz at all.

Deterministic ids

Every extractor assigns its own doc_id, and uid ("{source}:{doc_id}") is load-bearing downstream: the curation stage (slm4ie/data/curate/convert.py) requires a non-empty uid and uses it as the datatrove document id, raising KeyError if it is missing. So an id must be unique and reproducible for a given snapshot.

Most extractors carry an id straight out of the source (<doc id=...> for macocu, # newdoc id for conllu, xml:id for tei, a named field for jsonl / json / coleslaw). The two that have no source id — text and huggingfacederive one instead:

  • text keys off the file path plus the block's position in that file (sl:000123).
  • huggingface prefers a natural key column and otherwise keys off the config, split, and row index (sl:train:00000042).

Both derivations depend only on the input, not on how the work was scheduled — in particular, text produces identical ids under the serial and sharded writers, because sharding splits the file list on whole-file boundaries and preserves per-file block order.

The orchestrator still holds a positional fallback for a document that arrives with doc_id unset — idx-{index:014d} on the serial path, or a shard-namespaced idx-{shard:05d}-{local:010d} when sharded. No shipped extractor reaches it. It is worth knowing that the two schemes disagree: a future id-less extractor would get worker-count-dependent ids, so a new extractor should assign its own id rather than lean on the fallback.

Ids are snapshot-stable, not content-stable

A derived id points at a position in a specific snapshot of the source. If upstream re-releases a corpus with rows inserted or files renamed, re-extraction assigns different ids to the same text. This is fine for the six web corpora that rely on derivation — all are role: pretrain and never feed the uid-hashed train/val/test split logic in to_spans / to_sentiment. A content hash was considered and rejected: exact-duplicate text is common in web corpora and would collide.

json and jsonl share field-mapping

Both honour the same metadata: knobs — text_field, id_field, and the metadata_fields whitelist — so field names are configurable on either. They differ only in container shape (json reads a top-level array or single object, jsonl reads one object per line) and in that jsonl alone parses nested paragraphs into annotations. Pick by input layout, not by config capability.

Sharding is narrower than it looks

Only text, macocu, conllu, and tei subclass FileBasedExtractor. jsonl, json, coleslaw, and huggingface are plain BaseExtractors and always run single-pass, regardless of --max-workers.

Adding an extractor

Subclass BaseExtractor (or FileBasedExtractor, to get sharding for free by splitting enumeration from parsing), then call register_extractor("<name>", <Class>) at module scope and import the module in slm4ie/data/processing.py so registration fires. Point a new extract.yaml entry at the registry name.