# .l10n-sync — pt-BR localization pipeline

Semi-automatic pipeline for delivering and maintaining the Brazilian Portuguese
(`pt-BR`) localization of the Ripple Payments Direct 2.0 documentation
(`products/payments-direct-2/@v2026.03/` only). Standing rules: `CLAUDE.md` at the
repo root. Bootstrap findings: `BOOTSTRAP-REPORT.md` in this directory.

## Contents

| Path | Role |
|  --- | --- |
| `scope.json` | The in-scope source file list. Grow it as coverage expands. |
| `generate-manifest.mjs` | Builds `manifest.json`; `--check` mode is the drift check. Also exports `splitSections()` so other scripts share one definition of a section boundary. |
| `selftest.mjs` | Regression checks for the sectioning invariants (DOC-6071). Synthetic fixtures, so it does not break when docs content changes. Exit 1 on failure. |
| `tm-lookup.mjs` | Given the pages you are about to translate, prints only the memory entries whose English appears in them. Exact matching, on the same segmentation the harvest used. `--stats` for coverage only, `--json` for a JSONL subset. |
| `harvest-tm.mjs` | Extracts approved EN/pt-BR segment pairs into `tm.jsonl`. Requires `--mr` and `--date` for provenance; `--dry-run` reports without writing. Run it only over reviewer-approved content: the script cannot tell approved pages from unapproved ones. |
| `manifest.json` | Generated. Per-file, per-section sha256 hashes of the English source at translation time. |
| `tm.jsonl` | Translation memory. **Append-only**: one JSON object per line, `{"en": ..., "ptBR": ..., "source": "<file>#<sectionId>", "approvedIn": "<MR ref>", "date": "YYYY-MM-DD"}`. Populated on native-reviewer approval of an MR on its draft branch — approval, not merge, is the trigger (`CLAUDE.md` §8.3); never before that approval lands. Consulted on every translation pass. Corrections are appended rather than overwritten and resolved last-wins on lookup; see *Correcting an approved record*. |
| `termbase/` | `ptbr_payments_termbase_seed.xlsx` (human-edited source of record) + generated CSVs (`termbase.csv`, `do-not-translate.csv`, `idiom-traps.csv`). Regenerate CSVs with `python3 .l10n-sync/termbase/export_csv.py`. |
| `BOOTSTRAP-REPORT.md` | Session-1 bootstrap report (mechanics verification, inventory, decisions). |


## The manifest

`manifest.json` maps each in-scope English source file to its mirrored pt-BR target
(`@l10n/pt-BR/<source path>`, version folder included; mechanics empirically verified
2026-07-21) and records a sha256 hash per heading-delimited section. `.md` files are
split on ATX headings (front matter and pre-heading preamble are their own sections);
the OpenAPI `.yml` is a single section (spec updates are bulk events).

### Section identity (DOC-6071)

A section's id is the slugified path of its ancestor headings, for example
`create-a-payment/step-3-create-a-payment-with-the-v3-payments-operation/endpoint`.
Identity is structural rather than positional, so inserting or moving a heading leaves
every other section's id untouched. A `~n` suffix disambiguates a genuine path
collision; the current corpus contains none.

This replaced an earlier `<index>:<slug>` scheme under which one inserted heading
renumbered every section below it. Measured against real drift (16 docs commits between
the pilot and 2026-09-11), the old scheme reported 30 stale sections of which 17 had
byte-identical English. The path scheme reports 12 stale and no false positives.

Two consequences worth knowing:

- **Renaming a heading changes its id**, so a rename reads as one section removed and
another added rather than as a modification. That is usually the right outcome: a
renamed English heading needs its pt-BR heading retranslated anyway.
- **Trailing blank lines and horizontal rules are excluded from the hash.** A separator
sits between sections rather than belonging to the one above it, so adding `---` ahead
of a new section no longer marks the preceding section stale.


Regenerate after any translation lands: `node .l10n-sync/generate-manifest.mjs`

## The sync task

On a "sync"/"update" request (run by the AI pipeline, output is always an MR):

1. Run `node .l10n-sync/generate-manifest.mjs --check` to list stale sections.
2. Treat a stale flag as a signal to verify, not an instruction to re-translate:
confirm the English actually changed before touching the pt-BR. Sections whose
English is unchanged keep their approved translation byte-identical.
3. For light English edits, minimally revise the previous approved pt-BR text
(consult `tm.jsonl`) — never re-translate from scratch.
4. New source pages in scope → add to `scope.json` and translate. Orphaned pt-BR
files (source deleted/moved) → **flag in the MR, never delete**.
5. Rewrite `#fragment` links inside translated files to match translated heading
slugs (anchor policy, CLAUDE.md §5.3), and run a fragment-link check.
6. Regenerate `manifest.json`; open an MR per CLAUDE.md §3 containing only what
changed.


## Translation memory workflow

After a native reviewer approves a localization MR on its draft branch, harvest the
approved segment pairs (including reviewer corrections from MR discussions) into
`tm.jsonl`:

```
node .l10n-sync/harvest-tm.mjs --dry-run              # preview
node .l10n-sync/harvest-tm.mjs --mr "!1359" --date 2026-08-12
```

### Consulting the memory

`CLAUDE.md` §8.2c and §8.3 require the memory to be consulted on every pass, but it grows
with coverage: ~41k tokens at 396 records, projecting to roughly 5,400 records once the
full docs set is translated. Reading it whole per page does not scale, so narrow it first:

```
node .l10n-sync/tm-lookup.mjs <the pages you are about to translate>
node .l10n-sync/tm-lookup.mjs products/payments-direct-2/@v2026.03 --stats
```

Matching is exact. A near-match is a judgment call that belongs to the translator rather
than to a lookup table.

Two things to expect from the output. Most matches are a small number of very frequent
short segments: on `financial-instruments.md`, 296 matches come from 14 distinct entries,
and `Yes` alone accounts for 115 of them. The value there is consistency rather than saved
effort, since rendering `Yes` two ways across one page is a defect a reviewer has to catch
repeatedly. Entries of two words or fewer are marked `[short]` because they can be
context-sensitive; confirm the sense before reusing one.

**Harvest after each batch, not once at the end.** Reuse only exists for content already
harvested, so a batch translated before the previous batch is harvested gets none of it.

### How pairing works

Segments are line-level. Sections are paired by position rather than by id, because
pt-BR headings are translated and the ids therefore differ between the two trees;
position is reliable only because §5 requires structure to be preserved, so the script
verifies section counts, heading levels, and per-line structural types before pairing
and refuses to harvest anything that has diverged. Pairs whose English and Portuguese
are identical are skipped: link text pointing at untranslated pages is deliberately left
in English, and recording it would teach the memory to translate those phrases as
themselves. Re-running adds only records that are new by the composite key
`en >> ptBR`, so writes are idempotent; a corrected translation for unchanged English is
new by that key and is appended, and the lookup resolves it last-wins (see *Correcting an
approved record* below). Approval is the trigger, not merge: the
`l10n/pt-br/*` branches do not merge
to the default branch (see `CLAUDE.md` §8.3). Reviewer corrections land via MR review
changes only — pt-BR files are generated artifacts and are never hand-edited outside MR
review (`CLAUDE.md` §8.6).

### Correcting an approved record

`tm.jsonl` is append-only, so a reviewer correction to Portuguese whose English has not
changed cannot overwrite the record it replaces. `harvest-tm.mjs` dedups on the `en >> ptBR` pair, so the correction is appended as a second record for the same English. The
resolution rule that makes this work is in `tm-lookup.mjs`:

**Last record wins.** `resolveMemory()` keeps the last record for each English segment
rather than the first. Where the memory holds more than one Portuguese for a segment, the
lookup says so rather than picking quietly:

```
Translation-memory lookup: 2 file(s), 396 records in memory
  multiple wordings   : 10 of these have more than one Portuguese in memory; the last is served
```

That count is **distinct segments, not displaced records** (three records for one segment
is one entry the translator has to think about), and it is **scoped to what the lookup
matched**, not to the whole memory. An ambiguity in a part of the memory these pages do
not touch is not the translator's problem, and counting it would read as "some of the
entries below are ambiguous" when none of them are. In `--json` mode the equivalent notice
goes to stderr so it cannot corrupt the JSONL on stdout.

**So correcting approved wording is just a re-harvest.** Fix the pt-BR through MR review,
merge, re-harvest. No pruning, and the displaced record keeps its `approvedIn` and `date`
provenance in the file.

#### Known limitations

Read these before trusting the rule. They are why the output says "multiple wordings"
rather than "corrected".

**The rule is byte order, not approval order.** Within a single harvest run, records are
appended in manifest order. If the same English legitimately takes different Portuguese in
two files, the winner is whichever file the manifest lists later. That is arbitrary.
First-wins was equally arbitrary in the other direction, so this is not a regression, but
the rule must not be described as "later approved". **`tm-lookup` cannot tell a reviewer
correction from two files disagreeing**, which is why it does not claim to.

Measured 2026-09-14 on the 5-file pilot set: 8 English segments appear in more than one
in-scope file, and 0 of them have divergent Portuguese. That is the whole collision surface
today. At full coverage the lookup reports 142 files and 12,806 segments, and the segments
most likely to diverge are short context-sensitive ones, which is what the `[short]` tag
already warns about. **Expect this to stop being theoretical as coverage grows.** A real
fix means context-aware memory, which is a different piece of work.

**`date` and `approvedIn` are not consulted.** A record appended later with an earlier date
still wins. Ordering therefore depends on how `tm.jsonl` was written, including how git
ordered it if two branches both appended and were merged. Do not reorder or hand-edit the
file; `harvest-tm.mjs` appending is the only writer the rule assumes.

**The fix lives in the tool, not in the data.** `CLAUDE.md` §8.3 says to consult the memory
on every pass; it does not mandate `tm-lookup`. A `grep` on `tm.jsonl` still returns every
record for a segment, superseded ones included. Anything that reads the file directly,
including a future agent session, still sees the ambiguity. Use `tm-lookup`.

> **History.** Until 2026-09-14 the lookup resolved *first*-wins, which meant a correction
was written to `tm.jsonl` and then never served: the stale wording kept being offered as
approved, with no warning, and the only symptom was a digest header reporting fewer
records than the file had lines. Found when 10 approved records had their Portuguese
corrected on `terminology-corrections`. `selftest.mjs` now covers the resolution rule;
a regression there is silent and expensive, exactly like the sectioning invariants.


## Quality gates: run both before opening a translation MR

Two checks, deliberately separate. Structure cannot see meaning; terminology cannot see
structure. Each exists because something got through without it.

```
node .l10n-sync/qc-structure.mjs      # gate 1: does pt-BR mirror the source?
node .l10n-sync/qc-terminology.mjs    # gate 2: termbase, register, pt-BR variant
```

Both exit non-zero on findings and accept an optional path substring to narrow the run.

**qc-structure** compares section counts, per-section line counts and per-line structural
kind, pipe-table cell counts, Markdoc token parity, fenced-block contents, `$env` vars,
code identifiers, trailing newline, and whether every fragment link resolves.

Two subtleties it encodes. Comparison is **per section**, not whole-file, because that is
how `harvest-tm` pairs and because blank lines between sections legitimately differ.
Fenced blocks are compared **ignoring comment lines**, because `CLAUDE.md` §4 allows
translating a code comment and nothing else inside a fence.

**qc-terminology** checks Do-Not-Translate parity, Verified termbase entries that appear in
the source but have no trace in the translation, retired renderings, `CLAUDE.md` §10
register rules, and pt-PT surface forms.

It reads **prose only**: fenced blocks and inline code are stripped first, so
`invoiceNumber=INV-2025-0615` is not reported as an untranslated "invoice". pt-PT findings
come in two grades: `[pt-PT]` is wrong in any sense, `[pt-PT?]` is wrong only in a specific
sense and needs a human to look — "transferir" is right for moving money and wrong for
"download".

Findings are advisory where the termbase itself is ambiguous. The point is to put the
question in front of a person, not to rewrite text automatically.

## Drift check (GitLab CI)

Not yet wired into `.gitlab-ci.yml` (currently the stock template; no docs jobs run
there today). Proposed addition to `.gitlab-ci.yml`, not landed. Landing requires the
DOC-6063 maintenance demonstration, which proves the manifest and drift mechanics end to
end, and an agreed route to the default branch, which does not currently exist for this
work. Until it lands, this draft is the reference copy and stays here:

```yaml
# Guards the sectioning invariants (DOC-6071). A regression here is a code
# defect, not a content signal, so it fails the pipeline rather than warning.
l10n-selftest:
  stage: test
  image: node:22-alpine
  rules:
    - if: $CI_PIPELINE_SOURCE == "merge_request_event"
      changes:
        - ".l10n-sync/**/*"
  script:
    - node .l10n-sync/selftest.mjs

# Reports when English has moved without the pt-BR being resynced. Warn-only
# during the pilot; remove allow_failure to enforce.
l10n-drift-check:
  stage: test
  image: node:22-alpine
  rules:
    - if: $CI_PIPELINE_SOURCE == "merge_request_event"
      changes:
        - "products/payments-direct-2/@v2026.03/**/*"
        - "@l10n/**/*"
        - ".l10n-sync/**/*"
  script:
    - node .l10n-sync/generate-manifest.mjs --check
  allow_failure: true   # warn during pilot; drop this line to enforce later
```

The two are separate jobs on purpose. Drift is informational while the pilot is
in progress, so it warns. A broken sectionizer would silently put approved
Portuguese back in a reviewer's queue, so it fails.

## Version carry-forward (planned)

`@v2026.04` exists but is config-ignored (release delayed pending an engineering fix;
Quotes V3, target mid-August). When localization moves to it: seed the new version's
pt-BR tree by matching `@v2026.04` sections against `manifest.json` hashes and
`tm.jsonl` — unchanged sections carry forward verbatim, changed ones go through the
normal sync flow. Expect high reuse (the trees are structurally near-identical).
Do not translate `@v2026.04` until explicitly instructed.