# Draft: l10n(pipeline): resolve the translation memory last-wins so reviewer corrections take effect [DOC-6101]

## 1. What this is

A pipeline correctness fix. Until now `tm-lookup.mjs` resolved the translation memory
**first**-wins, which meant a reviewer correction to already-harvested Portuguese was
written to `tm.jsonl` and then never served. The stale wording kept being offered to
translators as approved, silently.

- Branch: `l10n/pt-br/tm-resolution`, based on `l10n/pt-br/bootstrap` at `75cdc8fd9`.
Neither merges to the default branch.
- Jira: DOC-6101, under epic DOC-6059.
https://ripplelabs.atlassian.net/browse/DOC-6101
- No translation content in this MR. English source is unmodified.
- Found while validating the remediation described in the pt-BR handoff §3.2, which
prescribed a plain re-harvest. That remediation was written without being run, and
running it left all 10 affected records still being served.


## 2. Scope of this MR

- `tm-lookup.mjs`: resolution changed to last-wins; `resolveMemory()` factored out as a
pure, exported function; superseded records counted and reported.
- `selftest.mjs`: five checks covering the resolution rule.
- `.l10n-sync/README.md`: the behaviour documented, and the previous idempotence wording
corrected.


Not in this MR: any translated content, any change to `tm.jsonl`, `manifest.json`,
`scope.json`, the termbase, `redocly.yaml`, `.vale.ini`, or CI.

## 3. Files changed (CLAUDE.md §3: source files / sections changed)

No source files are translated by this MR, so the template's source table does not apply.
Pipeline files changed:

| File | Change | Why |
|  --- | --- | --- |
| `.l10n-sync/tm-lookup.mjs` | `loadMemory()` split into pure `resolveMemory(text)` plus a thin disk wrapper; resolution is last-wins; segments with more than one Portuguese are counted and reported in the digest header and, in `--json` mode, on stderr, scoped to the entries the lookup matched | The read-side half of the defect. Resolving on `en` first-wins made corrections unreachable |
| `.l10n-sync/selftest.mjs` | Five checks under a new `translation-memory resolution` group | The failure is silent, like the sectioning invariants the self-test already guards |
| `.l10n-sync/README.md` | New *Correcting an approved record* section with a **Known limitations** subsection; idempotence sentence and `tm.jsonl` table row corrected | The README said re-running the harvest was enough. It was not |


## 4. English source reference (CLAUDE.md §3: the English diff or link)

n/a. No English is referenced or translated. Base ref `75cdc8fd9`.

## 5. Mirrored-OpenAPI render check (Phase 2, gating)

n/a. No spec content in this MR.

## 6. Structural preservation (CLAUDE.md §5)

n/a. No files under `@l10n/pt-BR/` are touched.

## 7. Terminology and termbase (CLAUDE.md §3: new termbase entries proposed)

No termbase entries proposed and no termbase files changed.

Relevant to terminology work, though: this fix is what makes the
`l10n/pt-br/terminology-corrections` re-harvest behave correctly. Ten approved records on
that branch have corrected Portuguese against unchanged English, which is exactly the case
that was broken. See §15 for the merge-order consequence.

## 8. Reviewer attention items (CLAUDE.md §3)

1. **The last-wins rule is byte order, not approval order.** An earlier revision of this MR
justified it as "later in an append-only file means later approved". That is wrong, and
an adversarial pass on this branch found it before review did. Within a single harvest
run, `harvest-tm.mjs` appends records in **manifest order**, so if the same English
legitimately takes different Portuguese in two files, the winner is whichever file the
manifest lists later. Arbitrary. First-wins was equally arbitrary in the other
direction, so this is not a regression, but the justification was overstated and the
code comment, README, and DOC-6101 have been corrected. **The rule is defensible; the
original reasoning for it was not.** Worth a second opinion on whether byte order is an
acceptable basis at all.
2. **The tool cannot tell a correction from a collision.** Both look identical in an
append-only file. The digest therefore reports `multiple wordings ... the last is served` rather than claiming anything was corrected. An earlier revision did claim it,
which would have been false in the collision case, and would have manufactured
confidence in exactly the way `CLAUDE.md`-adjacent governance keeps getting bitten by.
Measured on the pilot set: **8 English segments appear in more than one in-scope file,
0 with divergent Portuguese.** That is today's whole collision surface. At full coverage
the lookup sees 142 files and 12,806 segments, so expect this to stop being theoretical.
3. **`date` and `approvedIn` are not consulted.** A record appended later with an earlier
date still wins. Ordering depends on how `tm.jsonl` was written, including how git
ordered it if two branches both appended and were merged, which has already happened on
this project (`tm-harvest` into `bootstrap`). Documented as a limitation rather than
fixed. Push back if you think it should be enforced.
4. **The fix is in the tool, not the data.** `CLAUDE.md` §8.3 says consult the memory on
every pass and does not mandate `tm-lookup`. A `grep` on `tm.jsonl` still returns both
records. The ambiguity is mitigated behind one tool, not eliminated.
5. **Rejected alternative.** A `supersedes` field written by `harvest-tm` and honoured by
`tm-lookup` would carry better provenance, but needs a migration for the existing 396
records and buys nothing that last-wins does not, since displaced records stay in the
file either way. Say so if you disagree.


## 9. English-screenshot gap log (CLAUDE.md §5.5)

n/a.

## 10. Source defects found (CLAUDE.md §9: reported, not fixed)

None in the English source. The defect fixed here is in our own pipeline, not in docs-team
content, so §9's report-do-not-fix rule does not apply.

## 11. Config changes (preview only; do not enable in production here)

None.

## 12. Translation memory

`tm.jsonl` is **unchanged**: 396 records, not touched by this MR. The change is read-side
only. No harvest was run against the repository.

Consequence for the workflow: correcting approved wording no longer needs `tm.jsonl` to be
pruned by hand. Fix the Portuguese through MR review, then re-harvest. The superseded
record stays in the file with its `approvedIn` and `date` provenance intact, and the
append-only invariant in `CLAUDE.md` §8.3 is preserved.

## 13. Pipeline / manifest

- `manifest.json` unchanged: 5 files, 69 sections.
- `generate-manifest.mjs --check`: **green**, `manifest up to date`.
- `selftest.mjs`: **all checks passed**, the five new ones alongside the existing
sectioning and segment-extraction groups.
- `harvest-tm.mjs --dry-run`: **0 new records**, unchanged.


Verification of the change itself:

- **No-op on current data.** `bootstrap` holds 396 records across 396 distinct English
keys, with zero conflicting pairs. Patched output on the two pilot pages is
byte-identical to pre-fix output, 1,133 lines, compared with `diff`.
- **Fixes the real case.** Simulated on the post-merge tree for
`terminology-corrections` (a fast-forward, so that tree is byte-identical to the branch):
after a plain re-harvest with no pruning, `tm-lookup` serves 0 superseded and 10
corrected, and reports `multiple wordings : 10 of these have more than one Portuguese in memory; the last is served`.
- **The tests can fail.** Reverting the resolution rule to first-wins and leaving the rest
of the change in place fails 3 of the 5 new checks. The two that still pass are the ones
that should: an unchanged re-harvest supersedes nothing under either rule, and a
malformed line is ignored by both.


## 14. Review gate

Do not self-approve, mark ready, merge, or resolve discussions.

This is pipeline code rather than Portuguese, so it does not need native review to be
judged correct. It does change what translators are served, so Luis should be aware of it
even if he does not review the diff.

Reviewer checklist:

- [ ] The last-wins premise holds for how `tm.jsonl` is actually written (§8.1).
- [ ] The arbitrary-resolution trade-off is acceptable, or a better rule is proposed (§8.2).
- [ ] The rejected `supersedes` alternative was rejected for the right reasons (§8.3).
- [ ] `resolveMemory()` being exported for the self-test is an acceptable seam.
- [ ] The README now describes what the code does.


## 15. Open items and human calls

1. **Merge order.** If `terminology-corrections` is re-harvested before this lands, the
harvest still needs the prune-first workaround and someone has to remember. Landing this
first removes that step. This MR does not depend on that one.
2. **Should `harvest-tm.mjs` say something too?** It currently reports `new records: N`
without distinguishing a brand-new segment from a correction to an existing one. Naming
that difference at harvest time would catch an unintended correction earlier. Deliberately
not done here to keep the change small.
3. **The handoff and the README both carry the history** of the first-wins behaviour, so a
future reader does not re-derive it. Remove those notes only when they stop being useful.