# Online encoding migration (design note) **Status:** design **Branch:** `uw-migration` **Scope (MWP):** change **key/value encoding** of an existing **set** table by copying into a new RocksDB column family while the table remains online. Changing Mnesia `type` (`set` ↔ `ordered_set`) is **out of scope** for the first cut. For `mrdb` users, what matters for range seeks and key order is the **rocksdb key encoding**, not the Mnesia metadata type. Aligning Mnesia type via a custom schema transaction is possible later and needs separate design. --- ## 1. Motivation Some tables were created with encodings that are awkward for production use. Example: integer primary keys with `{term, {object, term}}` encoding. The Erlang external term format preserves sort order for positive integers only up to 32-bit max; beyond that keys use `SMALL_BIG_EXT` and **byte order ≠ numeric order**. RocksDB iterators and seeks then mis-order heights. Historically apps worked around this by creating a second table (`gmmp_tallies` → `gmmp_tallies2`) and migrating in application code. That is painful and easy to get wrong. This feature moves **encoding migration into the backend**: dual-write to a new column family (CF), copy old → new, then flip the live `db_ref` and drop the old CF. --- ## 2. Goals and non-goals ### Goals (MWP) - Online migration of **encoding** for **set** tables (semantics `set`). - Table remains readable and writable during migration. - Safe under concurrent inserts/updates/deletes. - Resume after process/node restart (without a durable RocksDB snapshot). - Indexes (ordered) continue to work: updated on live writes as today; index CFs are **not** dual-copied for encoding-only main-table migrate. - Explicit start/complete as admin/schema-level operations (not silent). ### Non-goals (MWP) - **Bag** tables (supported for completeness; poor fit for rocksdb dual-write). - Changing Mnesia **type** (`set` / `ordered_set` / `bag`) in this pass. - Rebuilding or re-encoding **index** column families. - Multi-node coordinated migration (single backend instance first). - Retaining a RocksDB snapshot across process restart. - Record shape / arity / attribute renames (use `mnesia:transform_table` or app logic). --- ## 3. Existing building blocks | Piece | Role today | |---|---| | `persistent_term` meta map | `Name → db_ref()` via `put_pt` / `get_ref` / `ensure_ref` | | CF create/drop | `create_column_family` / `drop_column_family` | | Offline standalone→CF migrate | `migrate_standalone/2`: chunk `select`, insert new, delete old, `put_pt` | | Snapshots | `mrdb:snapshot/1`, `release_snapshot/1` (DB-instance scoped) | | Hot write path | `insert_` / `delete_` encode with ref’s `encoding` | The new feature is an **online, same-DB CF** migrate with dual-write, not a reimplementation of standalone migration. --- ## 4. Design overview ### 4.0 Column-family naming (versions) RocksDB CF names must be **unique** within a DB. The existing map is: | Logical resource | Physical CF name | |---|---| | data table `Tab` | `"{d, Tab}"` (gen 0) | | index | `"{i, Tab, I}"` | | retainer | `"{r, Tab, R}"` | | admin | `"{a, Alias}"` | That mapping has **no version**. Encoding migration needs the old and new data CFs to coexist until cutover, so data CFs are versioned: | Generation | Physical CF name | |---|---| | 0 (legacy / first create) | `"{d, Tab}"` | | N ≥ 1 | `"{d, Tab, N}"` | - `tab_to_cf_name(Tab)` still produces gen 0 for brand-new tables. - Migration target = `data_cf_name(Tab, LiveGen + 1)`. - Live gen is stored on the `db_ref` as `cf_gen` / `cf_name` and durably as admin info `cf_gen` for the table. - `cf_name_to_tab/2` maps **any** data gen back to logical `Tab`. - On open, if several data gens exist for `Tab`, prefer the highest gen for which we have a handle; refine with durable `cf_gen` when present (interrupted migration may leave an incomplete higher gen — recovery should prefer the recorded live gen once that path is fully wired). **Cutover does not recreate a “canonical” CF name.** The versioned target CF *is* the new live CF; the previous generation is dropped. No second full copy. ``` ┌─────────────────────────────────────┐ │ Old CF gen G (reads; encoding E_old)│ │ name: {d,Tab} or {d,Tab,G} │ └──────────────┬──────────────────────┘ │ dual-write │ (live updates) ▼ ┌─────────────────────────────────────┐ │ New CF gen G+1 (encoding E_new) │ │ name: {d,Tab,G+1} │ │ + live dual-writes │ │ + migrator copy (skip if exists) │ └─────────────────────────────────────┘ │ ┌──────────────┴──────────────────────┐ │ Admin info │ │ encoding_migration meta, cf_gen │ └─────────────────────────────────────┘ ``` ### 4.1 Phases | Phase | Reads | Writes | Notes | |---|---|---|---| | `idle` | current CF | current CF | Normal operation | | `dual_write` / `copying` | **old** CF only | old **and** new (re-encode) | Migrator walks old → new | | `copy_done` | old | dual-write still on | Optional verify | | `cutover` | **new** | **new** only | PT flip; then drop old + scratch | ### 4.2 Ref shape during migration Live ref in persistent_term (conceptual): ```erlang OldRef#{ migration => NewRef, %% target CF ref (encoding E_new) migration_epoch => pos_integer(), migration_phase => copying | copy_done } ``` - **Reads / select / iterators:** use the outer (old) ref only until cutover. - **Inserts / deletes / merges (set):** apply to old, then if `migration` is set, re-encode and apply to `NewRef`. - **Indexes:** `update_index` runs **once**, as today, using the live main ref (old CF) for `read(R, Key)`. Index CF layout is independent of main encoding (index keys are `{IxVal, Key}` under the **index** ref’s encoding). ### 4.3 Migrator install rule For each key `K` observed on the migration walk: ``` if not present in live old(K) → skip %% deleted after snapshot / walk else if present in new(K) → skip %% dual-write already installed latest else put new(live_value(K)) %% re-encode with NewRef ``` Rationale: - **Skip if new has key** — dual-write wins; migrator must not clobber newer data. - **Live re-check of old** — avoids **delete resurrection** when iterating a snapshot that still contains `K` after a dual-delete removed it from both CFs. - Prefer **live** value over snapshot-only value when installing. ### 4.4 Snapshot (best-effort) While the migrator process lives: 1. `Snapshot = rocksdb:snapshot(DbRef)` (same DB as old/new CFs). 2. Iterate **old** CF with `{snapshot, Snapshot}` read options from `cursor`. 3. Apply the install rule above (always re-check live old + new). On migrator crash/restart: snapshot is gone. Resume from scratch **cursor** on **live** old CF with the same install rule. Correctness does not depend on the snapshot; the snapshot only reduces iterator churn. ### 4.5 Scratch metadata Either a dedicated scratch CF or admin-info keys (admin CF is enough for MWP). Minimum fields: ```erlang #{ table => tabname() , phase => dual_write | copying | copy_done | cutover , target => #{encoding := encoding()} %% validated , cursor => '$first' | Key :: term() %% last successfully processed logical key , epoch => pos_integer() , started_at => millisecond() } ``` - Cursor is the **logical** key (Erlang term), not the encoded rocksdb key. - Resume seek on old CF uses **old** encoding: `encode_key(Cursor, OldRef)`. ### 4.6 Cutover 1. Ensure `phase = copy_done` (and optional verification). 2. **Barrier:** no in-flight activities that still hold a pre-migration ref map (see §5.2), or epoch re-resolution (see below). 3. Schema/admin transaction: - `put_pt(Tab, NewRef)` without `migration` (NewRef is gen G+1 CF) - persist live `cf_gen = G+1` in admin info - drop **old** generation CF only - clear migration meta 4. Subsequent `get_ref(Tab)` returns new encoding and `cf_gen = G+1`. No rename and no second bulk copy: the migration target CF is permanent. ### 4.7 Mnesia type deferred For `mrdb` callers, ordered seeks and range logic depend on **key encoding** (e.g. `sext` vs `term`), not on `mnesia:table_info(Tab, type)`. Changing Mnesia type later may use a custom schema transaction so metadata matches reality; that must not be conflated with CF encoding migrate and needs extra care around `table_info`, load hooks, and any code that branches on type. --- ## 5. Concurrency and correctness ### 5.1 Dual-write before copy Order of start: 1. Create empty new CF (target encoding). 2. Persist migration meta + install `migration => NewRef` on live ref (**dual-write ON**). 3. Only then start the copy walk. If copy ran before dual-write, concurrent writes could land only on old and be missed or overwritten incorrectly. ### 5.2 Stale in-process `db_ref` maps `ensure_ref/1` short-circuits when given a map (activities/batches hold refs). After cutover, a process still holding the **old** map would write a CF that is about to be dropped. MWP options (pick at least one): 1. **Quiesce:** refuse cutover while table activity is non-zero / use a short exclusive period. 2. **Epoch:** bump `migration_epoch`; `ensure_ref(Map)` re-fetches from PT if map epoch ≠ current. 3. **Delay drop:** keep old CF read-only until epoch drained (refcounting). Recommended MWP: **epoch on ref + re-resolve when mismatch**, plus **delay drop of old CF** until a quiet period or explicit admin “finalize”. ### 5.3 Transactions and batches Prefer one RocksDB `WriteBatch` (same `db_ref`, two `cf_handle`s) so dual-write insert/delete is atomic across old and new under `as_batch`. Index updates stay in the same batch as today (`batch_if_index`). ### 5.4 Indexes and encoding independence Index CFs store `{IxVal, MainKey}` under the **index** encoding. Main encoding change does not require rewriting index keys. Live writes continue to maintain indexes against the logical object; after cutover, `index_read` still yields `MainKey` terms and reads the new main CF with the new encoding. **Do not** dual-write index CFs in MWP encoding migration. ### 5.5 Crash recovery On admin/backend restart: 1. Reload CF handles and PT from durable admin state. 2. If migration meta says `copying` / `dual_write` / `copy_done`: - Re-bind live ref with `migration => NewRef`. - Resume dual-write. - If not `copy_done`, restart migrator from `cursor` (live iterate). 3. Never drop old CF unless cutover fully committed. --- ## 6. API sketch (MWP) ```erlang %% Start (admin / schema-level). Validates encoding; rejects bag; rejects %% unsupported type changes. -spec migrate_encoding(alias(), tabname(), encoding()) -> ok | {error, term()}. migrate_encoding(Alias, Tab, NewEncoding) -> ... %% Optional: progress / status -spec migration_status(tabname()) -> idle | #{phase := atom(), cursor := term(), target := map()}. %% Complete cutover (if not automatic after copy_done + barrier) -spec finalize_migration(alias(), tabname()) -> ok | {error, term()}. ``` Reporting can mirror `migrate_standalone/3` (`Rpt` pid + progress messages). --- ## 7. Risks and open points | Risk | Mitigation | |---|---| | Delete resurrection under snapshot | Live re-check before install (§4.3) | | Stale activity refs after cutover | Epoch re-resolve + delayed CF drop (§5.2) | | Partial dual-write in batch | Multi-CF WriteBatch (§5.3) | | PT lost while admin meta says migrating | Recovery rebuilds PT from admin (§5.5) | | Huge tables | Chunked walk + cursor; progress reports | | bag / type / index encoding | Explicitly out of MWP | Open (later): - Automatic vs manual finalize after `copy_done`. - Optional full verification (count / sample) before cutover. - Mnesia type alignment schema transaction. - Bag and index-encoding migration. --- ## 8. Relation to application workarounds Application dual tables (`*2`) remain valid for record-shape changes. Backend encoding migration is aimed at cases like “same record, wrong key encoding” without app-visible table rename. Example target: migrate `gmmp_generations` from `{term,{object,term}}` to a sext (or other order-preserving) key encoding so height seeks stay valid past 2³² − 1 without a `generations2` table. --- # Implementation checklist Phased so each phase is testable. Bags and Mnesia type change stay excluded. ## Phase 0 — Spec freeze and hooks - [ ] Freeze MWP scope: **set + encoding only**; document rejections for bag / type. - [ ] Define durable migration meta schema (admin info keys vs scratch CF). - [ ] Define public API: `migrate_encoding/3`, `migration_status/1`, `finalize_migration/2` (names flexible). - [ ] Define progress report messages (align with `migrate_standalone` rpt). - [ ] List all write paths that must dual-write: insert, delete, delete_object, update_counter/merge (if encoding-compatible), clear_table policy (reject migrate while clearing / dual-clear both). ## Phase 1 — Ref + dual-write plumbing - [ ] Extend live `db_ref` with optional `migration`, `migration_epoch`, `migration_phase` (or keep phase only in admin meta). - [ ] `ensure_ref/1`: when map has stale epoch, re-fetch from PT (or document quiesce-only MWP if epoch deferred). - [ ] Implement `dual_put` / `dual_delete` helpers: old encoding + new encoding; same WriteBatch when `as_batch` active. - [ ] Wire dual-write into `insert_` / `delete_` / related set paths only. - [ ] Confirm `update_index` still runs once and reads from **old** main ref during migration; no index dual-write. - [ ] Unit tests: concurrent insert/delete while `migration` set updates both CFs; index remains consistent with logical data. ## Phase 2 — Create target CF + start migration - [ ] Validate `NewEncoding` via `mnesia_rocksdb_lib:check_encoding/2`. - [ ] Reject bag semantics; reject no-op (same encoding); reject if migration already active. - [ ] Create empty new CF with target encoding (and consistent vsn/access_type). - [ ] Persist migration meta; install dual-write on live ref (schema/admin txn). - [ ] Test: after start, reads still see old data; writes appear on both CFs (inspect via raw/ref). ## Phase 3 — Copy walk + scratch cursor - [ ] Migrator process (supervised or admin-linked) with optional RocksDB snapshot. - [ ] Iterate old CF from cursor; for each object apply install rule (§4.3): - [ ] live old missing → skip - [ ] new present → skip - [ ] else put re-encoded live value to new - [ ] Update scratch cursor after each chunk (batch cursor updates). - [ ] Progress reporting every N keys. - [ ] On completion set `phase = copy_done`. - [ ] Tests: - [ ] empty table - [ ] populated table, no concurrent writes - [ ] concurrent overwrites (new has key → migrator skips; final value correct) - [ ] concurrent deletes (no resurrection) - [ ] restart migrator mid-way (cursor resume, no snapshot) ## Phase 4 — Recovery - [ ] On backend/admin start: detect incomplete migration from durable meta. - [ ] Re-install dual-write ref; restart migrator from cursor if needed. - [ ] Test kill -9 / admin restart mid-copy; complete and verify data. ## Phase 5 — Cutover + drop - [ ] Barrier policy implemented (epoch and/or quiesce and/or delayed drop). - [ ] `finalize_migration`: PT → NewRef only; clear migration fields. - [ ] Drop old CF; clear scratch/meta. - [ ] Test: post-cutover read/write/select/index_read; encoding is E_new. - [ ] Test: in-flight activity behaviour under chosen barrier policy. ## Phase 6 — Optional verify + polish - [ ] Optional key-count or sample verification under dual-write before cutover. - [ ] Metrics: keys copied, skipped (exists), skipped (deleted), duration. - [ ] Docs: user guide section + this design note linked from README. - [ ] Dialyzer / property tests for encode round-trip old→new. ## Phase 7 — Explicitly later (not MWP) - [ ] Mnesia `type` change via custom schema transaction. - [ ] Bag table migration. - [ ] Index CF encoding rebuild. - [ ] Multi-node / multi-replica orchestration. --- ## 9. Suggested first spike (days, not weeks) 1. Dual-write only (no migrator): force `migration => NewRef` in test; prove insert/delete hit both CFs; indexes still correct. 2. Migrator on quiescent table: full copy + cutover + drop. 3. Add concurrent writes/deletes + install rule. 4. Add cursor + restart. 5. Snapshot as optimization only. --- ## 10. Document history | Date | Note | |---|---| | 2026-08-03 | Initial design from gm_mining_pool / generations encoding discussion |