WIP migrate table encoding
This commit is contained in:
@@ -0,0 +1,440 @@
|
||||
# Online encoding migration (design note)
|
||||
|
||||
**Status:** design
|
||||
**Branch:** `uw-migration`
|
||||
**Scope (MWP):** change **key/value encoding** of an existing **set** table by
|
||||
copying into a new RocksDB column family while the table remains online.
|
||||
|
||||
Changing Mnesia `type` (`set` ↔ `ordered_set`) is **out of scope** for the first
|
||||
cut. For `mrdb` users, what matters for range seeks and key order is the
|
||||
**rocksdb key encoding**, not the Mnesia metadata type. Aligning Mnesia type
|
||||
via a custom schema transaction is possible later and needs separate design.
|
||||
|
||||
---
|
||||
|
||||
## 1. Motivation
|
||||
|
||||
Some tables were created with encodings that are awkward for production use.
|
||||
Example: integer primary keys with `{term, {object, term}}` encoding. The
|
||||
Erlang external term format preserves sort order for positive integers only
|
||||
up to 32-bit max; beyond that keys use `SMALL_BIG_EXT` and **byte order ≠
|
||||
numeric order**. RocksDB iterators and seeks then mis-order heights.
|
||||
|
||||
Historically apps worked around this by creating a second table
|
||||
(`gmmp_tallies` → `gmmp_tallies2`) and migrating in application code. That is
|
||||
painful and easy to get wrong.
|
||||
|
||||
This feature moves **encoding migration into the backend**: dual-write to a
|
||||
new column family (CF), copy old → new, then flip the live `db_ref` and drop
|
||||
the old CF.
|
||||
|
||||
---
|
||||
|
||||
## 2. Goals and non-goals
|
||||
|
||||
### Goals (MWP)
|
||||
|
||||
- Online migration of **encoding** for **set** tables (semantics `set`).
|
||||
- Table remains readable and writable during migration.
|
||||
- Safe under concurrent inserts/updates/deletes.
|
||||
- Resume after process/node restart (without a durable RocksDB snapshot).
|
||||
- Indexes (ordered) continue to work: updated on live writes as today;
|
||||
index CFs are **not** dual-copied for encoding-only main-table migrate.
|
||||
- Explicit start/complete as admin/schema-level operations (not silent).
|
||||
|
||||
### Non-goals (MWP)
|
||||
|
||||
- **Bag** tables (supported for completeness; poor fit for rocksdb dual-write).
|
||||
- Changing Mnesia **type** (`set` / `ordered_set` / `bag`) in this pass.
|
||||
- Rebuilding or re-encoding **index** column families.
|
||||
- Multi-node coordinated migration (single backend instance first).
|
||||
- Retaining a RocksDB snapshot across process restart.
|
||||
- Record shape / arity / attribute renames (use `mnesia:transform_table`
|
||||
or app logic).
|
||||
|
||||
---
|
||||
|
||||
## 3. Existing building blocks
|
||||
|
||||
| Piece | Role today |
|
||||
|---|---|
|
||||
| `persistent_term` meta map | `Name → db_ref()` via `put_pt` / `get_ref` / `ensure_ref` |
|
||||
| CF create/drop | `create_column_family` / `drop_column_family` |
|
||||
| Offline standalone→CF migrate | `migrate_standalone/2`: chunk `select`, insert new, delete old, `put_pt` |
|
||||
| Snapshots | `mrdb:snapshot/1`, `release_snapshot/1` (DB-instance scoped) |
|
||||
| Hot write path | `insert_` / `delete_` encode with ref’s `encoding` |
|
||||
|
||||
The new feature is an **online, same-DB CF** migrate with dual-write, not a
|
||||
reimplementation of standalone migration.
|
||||
|
||||
---
|
||||
|
||||
## 4. Design overview
|
||||
|
||||
### 4.0 Column-family naming (versions)
|
||||
|
||||
RocksDB CF names must be **unique** within a DB. The existing map is:
|
||||
|
||||
| Logical resource | Physical CF name |
|
||||
|---|---|
|
||||
| data table `Tab` | `"{d, Tab}"` (gen 0) |
|
||||
| index | `"{i, Tab, I}"` |
|
||||
| retainer | `"{r, Tab, R}"` |
|
||||
| admin | `"{a, Alias}"` |
|
||||
|
||||
That mapping has **no version**. Encoding migration needs the old and new data
|
||||
CFs to coexist until cutover, so data CFs are versioned:
|
||||
|
||||
| Generation | Physical CF name |
|
||||
|---|---|
|
||||
| 0 (legacy / first create) | `"{d, Tab}"` |
|
||||
| N ≥ 1 | `"{d, Tab, N}"` |
|
||||
|
||||
- `tab_to_cf_name(Tab)` still produces gen 0 for brand-new tables.
|
||||
- Migration target = `data_cf_name(Tab, LiveGen + 1)`.
|
||||
- Live gen is stored on the `db_ref` as `cf_gen` / `cf_name` and durably as
|
||||
admin info `cf_gen` for the table.
|
||||
- `cf_name_to_tab/2` maps **any** data gen back to logical `Tab`.
|
||||
- On open, if several data gens exist for `Tab`, prefer the highest gen for
|
||||
which we have a handle; refine with durable `cf_gen` when present (interrupted
|
||||
migration may leave an incomplete higher gen — recovery should prefer the
|
||||
recorded live gen once that path is fully wired).
|
||||
|
||||
**Cutover does not recreate a “canonical” CF name.** The versioned target CF
|
||||
*is* the new live CF; the previous generation is dropped. No second full copy.
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────┐
|
||||
│ Old CF gen G (reads; encoding E_old)│
|
||||
│ name: {d,Tab} or {d,Tab,G} │
|
||||
└──────────────┬──────────────────────┘
|
||||
│ dual-write
|
||||
│ (live updates)
|
||||
▼
|
||||
┌─────────────────────────────────────┐
|
||||
│ New CF gen G+1 (encoding E_new) │
|
||||
│ name: {d,Tab,G+1} │
|
||||
│ + live dual-writes │
|
||||
│ + migrator copy (skip if exists) │
|
||||
└─────────────────────────────────────┘
|
||||
│
|
||||
┌──────────────┴──────────────────────┐
|
||||
│ Admin info │
|
||||
│ encoding_migration meta, cf_gen │
|
||||
└─────────────────────────────────────┘
|
||||
```
|
||||
|
||||
### 4.1 Phases
|
||||
|
||||
| Phase | Reads | Writes | Notes |
|
||||
|---|---|---|---|
|
||||
| `idle` | current CF | current CF | Normal operation |
|
||||
| `dual_write` / `copying` | **old** CF only | old **and** new (re-encode) | Migrator walks old → new |
|
||||
| `copy_done` | old | dual-write still on | Optional verify |
|
||||
| `cutover` | **new** | **new** only | PT flip; then drop old + scratch |
|
||||
|
||||
### 4.2 Ref shape during migration
|
||||
|
||||
Live ref in persistent_term (conceptual):
|
||||
|
||||
```erlang
|
||||
OldRef#{
|
||||
migration => NewRef, %% target CF ref (encoding E_new)
|
||||
migration_epoch => pos_integer(),
|
||||
migration_phase => copying | copy_done
|
||||
}
|
||||
```
|
||||
|
||||
- **Reads / select / iterators:** use the outer (old) ref only until cutover.
|
||||
- **Inserts / deletes / merges (set):** apply to old, then if `migration`
|
||||
is set, re-encode and apply to `NewRef`.
|
||||
- **Indexes:** `update_index` runs **once**, as today, using the live main ref
|
||||
(old CF) for `read(R, Key)`. Index CF layout is independent of main
|
||||
encoding (index keys are `{IxVal, Key}` under the **index** ref’s encoding).
|
||||
|
||||
### 4.3 Migrator install rule
|
||||
|
||||
For each key `K` observed on the migration walk:
|
||||
|
||||
```
|
||||
if not present in live old(K) → skip %% deleted after snapshot / walk
|
||||
else if present in new(K) → skip %% dual-write already installed latest
|
||||
else put new(live_value(K)) %% re-encode with NewRef
|
||||
```
|
||||
|
||||
Rationale:
|
||||
|
||||
- **Skip if new has key** — dual-write wins; migrator must not clobber newer data.
|
||||
- **Live re-check of old** — avoids **delete resurrection** when iterating a
|
||||
snapshot that still contains `K` after a dual-delete removed it from both CFs.
|
||||
- Prefer **live** value over snapshot-only value when installing.
|
||||
|
||||
### 4.4 Snapshot (best-effort)
|
||||
|
||||
While the migrator process lives:
|
||||
|
||||
1. `Snapshot = rocksdb:snapshot(DbRef)` (same DB as old/new CFs).
|
||||
2. Iterate **old** CF with `{snapshot, Snapshot}` read options from `cursor`.
|
||||
3. Apply the install rule above (always re-check live old + new).
|
||||
|
||||
On migrator crash/restart: snapshot is gone. Resume from scratch **cursor**
|
||||
on **live** old CF with the same install rule. Correctness does not depend on
|
||||
the snapshot; the snapshot only reduces iterator churn.
|
||||
|
||||
### 4.5 Scratch metadata
|
||||
|
||||
Either a dedicated scratch CF or admin-info keys (admin CF is enough for MWP).
|
||||
|
||||
Minimum fields:
|
||||
|
||||
```erlang
|
||||
#{ table => tabname()
|
||||
, phase => dual_write | copying | copy_done | cutover
|
||||
, target => #{encoding := encoding()} %% validated
|
||||
, cursor => '$first' | Key :: term() %% last successfully processed logical key
|
||||
, epoch => pos_integer()
|
||||
, started_at => millisecond()
|
||||
}
|
||||
```
|
||||
|
||||
- Cursor is the **logical** key (Erlang term), not the encoded rocksdb key.
|
||||
- Resume seek on old CF uses **old** encoding: `encode_key(Cursor, OldRef)`.
|
||||
|
||||
### 4.6 Cutover
|
||||
|
||||
1. Ensure `phase = copy_done` (and optional verification).
|
||||
2. **Barrier:** no in-flight activities that still hold a pre-migration ref
|
||||
map (see §5.2), or epoch re-resolution (see below).
|
||||
3. Schema/admin transaction:
|
||||
- `put_pt(Tab, NewRef)` without `migration` (NewRef is gen G+1 CF)
|
||||
- persist live `cf_gen = G+1` in admin info
|
||||
- drop **old** generation CF only
|
||||
- clear migration meta
|
||||
4. Subsequent `get_ref(Tab)` returns new encoding and `cf_gen = G+1`.
|
||||
|
||||
No rename and no second bulk copy: the migration target CF is permanent.
|
||||
|
||||
### 4.7 Mnesia type deferred
|
||||
|
||||
For `mrdb` callers, ordered seeks and range logic depend on **key encoding**
|
||||
(e.g. `sext` vs `term`), not on `mnesia:table_info(Tab, type)`.
|
||||
|
||||
Changing Mnesia type later may use a custom schema transaction so metadata
|
||||
matches reality; that must not be conflated with CF encoding migrate and needs
|
||||
extra care around `table_info`, load hooks, and any code that branches on type.
|
||||
|
||||
---
|
||||
|
||||
## 5. Concurrency and correctness
|
||||
|
||||
### 5.1 Dual-write before copy
|
||||
|
||||
Order of start:
|
||||
|
||||
1. Create empty new CF (target encoding).
|
||||
2. Persist migration meta + install `migration => NewRef` on live ref (**dual-write ON**).
|
||||
3. Only then start the copy walk.
|
||||
|
||||
If copy ran before dual-write, concurrent writes could land only on old and be
|
||||
missed or overwritten incorrectly.
|
||||
|
||||
### 5.2 Stale in-process `db_ref` maps
|
||||
|
||||
`ensure_ref/1` short-circuits when given a map (activities/batches hold refs).
|
||||
After cutover, a process still holding the **old** map would write a CF that
|
||||
is about to be dropped.
|
||||
|
||||
MWP options (pick at least one):
|
||||
|
||||
1. **Quiesce:** refuse cutover while table activity is non-zero / use a
|
||||
short exclusive period.
|
||||
2. **Epoch:** bump `migration_epoch`; `ensure_ref(Map)` re-fetches from PT if
|
||||
map epoch ≠ current.
|
||||
3. **Delay drop:** keep old CF read-only until epoch drained (refcounting).
|
||||
|
||||
Recommended MWP: **epoch on ref + re-resolve when mismatch**, plus **delay
|
||||
drop of old CF** until a quiet period or explicit admin “finalize”.
|
||||
|
||||
### 5.3 Transactions and batches
|
||||
|
||||
Prefer one RocksDB `WriteBatch` (same `db_ref`, two `cf_handle`s) so dual-write
|
||||
insert/delete is atomic across old and new under `as_batch`.
|
||||
|
||||
Index updates stay in the same batch as today (`batch_if_index`).
|
||||
|
||||
### 5.4 Indexes and encoding independence
|
||||
|
||||
Index CFs store `{IxVal, MainKey}` under the **index** encoding. Main encoding
|
||||
change does not require rewriting index keys. Live writes continue to maintain
|
||||
indexes against the logical object; after cutover, `index_read` still yields
|
||||
`MainKey` terms and reads the new main CF with the new encoding.
|
||||
|
||||
**Do not** dual-write index CFs in MWP encoding migration.
|
||||
|
||||
### 5.5 Crash recovery
|
||||
|
||||
On admin/backend restart:
|
||||
|
||||
1. Reload CF handles and PT from durable admin state.
|
||||
2. If migration meta says `copying` / `dual_write` / `copy_done`:
|
||||
- Re-bind live ref with `migration => NewRef`.
|
||||
- Resume dual-write.
|
||||
- If not `copy_done`, restart migrator from `cursor` (live iterate).
|
||||
3. Never drop old CF unless cutover fully committed.
|
||||
|
||||
---
|
||||
|
||||
## 6. API sketch (MWP)
|
||||
|
||||
```erlang
|
||||
%% Start (admin / schema-level). Validates encoding; rejects bag; rejects
|
||||
%% unsupported type changes.
|
||||
-spec migrate_encoding(alias(), tabname(), encoding()) -> ok | {error, term()}.
|
||||
migrate_encoding(Alias, Tab, NewEncoding) -> ...
|
||||
|
||||
%% Optional: progress / status
|
||||
-spec migration_status(tabname()) ->
|
||||
idle | #{phase := atom(), cursor := term(), target := map()}.
|
||||
|
||||
%% Complete cutover (if not automatic after copy_done + barrier)
|
||||
-spec finalize_migration(alias(), tabname()) -> ok | {error, term()}.
|
||||
```
|
||||
|
||||
Reporting can mirror `migrate_standalone/3` (`Rpt` pid + progress messages).
|
||||
|
||||
---
|
||||
|
||||
## 7. Risks and open points
|
||||
|
||||
| Risk | Mitigation |
|
||||
|---|---|
|
||||
| Delete resurrection under snapshot | Live re-check before install (§4.3) |
|
||||
| Stale activity refs after cutover | Epoch re-resolve + delayed CF drop (§5.2) |
|
||||
| Partial dual-write in batch | Multi-CF WriteBatch (§5.3) |
|
||||
| PT lost while admin meta says migrating | Recovery rebuilds PT from admin (§5.5) |
|
||||
| Huge tables | Chunked walk + cursor; progress reports |
|
||||
| bag / type / index encoding | Explicitly out of MWP |
|
||||
|
||||
Open (later):
|
||||
|
||||
- Automatic vs manual finalize after `copy_done`.
|
||||
- Optional full verification (count / sample) before cutover.
|
||||
- Mnesia type alignment schema transaction.
|
||||
- Bag and index-encoding migration.
|
||||
|
||||
---
|
||||
|
||||
## 8. Relation to application workarounds
|
||||
|
||||
Application dual tables (`*2`) remain valid for record-shape changes. Backend
|
||||
encoding migration is aimed at cases like “same record, wrong key encoding”
|
||||
without app-visible table rename.
|
||||
|
||||
Example target: migrate `gmmp_generations` from `{term,{object,term}}` to a
|
||||
sext (or other order-preserving) key encoding so height seeks stay valid past
|
||||
2³² − 1 without a `generations2` table.
|
||||
|
||||
---
|
||||
|
||||
# Implementation checklist
|
||||
|
||||
Phased so each phase is testable. Bags and Mnesia type change stay excluded.
|
||||
|
||||
## Phase 0 — Spec freeze and hooks
|
||||
|
||||
- [ ] Freeze MWP scope: **set + encoding only**; document rejections for bag / type.
|
||||
- [ ] Define durable migration meta schema (admin info keys vs scratch CF).
|
||||
- [ ] Define public API: `migrate_encoding/3`, `migration_status/1`,
|
||||
`finalize_migration/2` (names flexible).
|
||||
- [ ] Define progress report messages (align with `migrate_standalone` rpt).
|
||||
- [ ] List all write paths that must dual-write: insert, delete, delete_object,
|
||||
update_counter/merge (if encoding-compatible), clear_table policy
|
||||
(reject migrate while clearing / dual-clear both).
|
||||
|
||||
## Phase 1 — Ref + dual-write plumbing
|
||||
|
||||
- [ ] Extend live `db_ref` with optional `migration`, `migration_epoch`,
|
||||
`migration_phase` (or keep phase only in admin meta).
|
||||
- [ ] `ensure_ref/1`: when map has stale epoch, re-fetch from PT (or document
|
||||
quiesce-only MWP if epoch deferred).
|
||||
- [ ] Implement `dual_put` / `dual_delete` helpers: old encoding + new encoding;
|
||||
same WriteBatch when `as_batch` active.
|
||||
- [ ] Wire dual-write into `insert_` / `delete_` / related set paths only.
|
||||
- [ ] Confirm `update_index` still runs once and reads from **old** main ref
|
||||
during migration; no index dual-write.
|
||||
- [ ] Unit tests: concurrent insert/delete while `migration` set updates both
|
||||
CFs; index remains consistent with logical data.
|
||||
|
||||
## Phase 2 — Create target CF + start migration
|
||||
|
||||
- [ ] Validate `NewEncoding` via `mnesia_rocksdb_lib:check_encoding/2`.
|
||||
- [ ] Reject bag semantics; reject no-op (same encoding); reject if migration
|
||||
already active.
|
||||
- [ ] Create empty new CF with target encoding (and consistent vsn/access_type).
|
||||
- [ ] Persist migration meta; install dual-write on live ref (schema/admin txn).
|
||||
- [ ] Test: after start, reads still see old data; writes appear on both CFs
|
||||
(inspect via raw/ref).
|
||||
|
||||
## Phase 3 — Copy walk + scratch cursor
|
||||
|
||||
- [ ] Migrator process (supervised or admin-linked) with optional RocksDB snapshot.
|
||||
- [ ] Iterate old CF from cursor; for each object apply install rule (§4.3):
|
||||
- [ ] live old missing → skip
|
||||
- [ ] new present → skip
|
||||
- [ ] else put re-encoded live value to new
|
||||
- [ ] Update scratch cursor after each chunk (batch cursor updates).
|
||||
- [ ] Progress reporting every N keys.
|
||||
- [ ] On completion set `phase = copy_done`.
|
||||
- [ ] Tests:
|
||||
- [ ] empty table
|
||||
- [ ] populated table, no concurrent writes
|
||||
- [ ] concurrent overwrites (new has key → migrator skips; final value correct)
|
||||
- [ ] concurrent deletes (no resurrection)
|
||||
- [ ] restart migrator mid-way (cursor resume, no snapshot)
|
||||
|
||||
## Phase 4 — Recovery
|
||||
|
||||
- [ ] On backend/admin start: detect incomplete migration from durable meta.
|
||||
- [ ] Re-install dual-write ref; restart migrator from cursor if needed.
|
||||
- [ ] Test kill -9 / admin restart mid-copy; complete and verify data.
|
||||
|
||||
## Phase 5 — Cutover + drop
|
||||
|
||||
- [ ] Barrier policy implemented (epoch and/or quiesce and/or delayed drop).
|
||||
- [ ] `finalize_migration`: PT → NewRef only; clear migration fields.
|
||||
- [ ] Drop old CF; clear scratch/meta.
|
||||
- [ ] Test: post-cutover read/write/select/index_read; encoding is E_new.
|
||||
- [ ] Test: in-flight activity behaviour under chosen barrier policy.
|
||||
|
||||
## Phase 6 — Optional verify + polish
|
||||
|
||||
- [ ] Optional key-count or sample verification under dual-write before cutover.
|
||||
- [ ] Metrics: keys copied, skipped (exists), skipped (deleted), duration.
|
||||
- [ ] Docs: user guide section + this design note linked from README.
|
||||
- [ ] Dialyzer / property tests for encode round-trip old→new.
|
||||
|
||||
## Phase 7 — Explicitly later (not MWP)
|
||||
|
||||
- [ ] Mnesia `type` change via custom schema transaction.
|
||||
- [ ] Bag table migration.
|
||||
- [ ] Index CF encoding rebuild.
|
||||
- [ ] Multi-node / multi-replica orchestration.
|
||||
|
||||
---
|
||||
|
||||
## 9. Suggested first spike (days, not weeks)
|
||||
|
||||
1. Dual-write only (no migrator): force `migration => NewRef` in test; prove
|
||||
insert/delete hit both CFs; indexes still correct.
|
||||
2. Migrator on quiescent table: full copy + cutover + drop.
|
||||
3. Add concurrent writes/deletes + install rule.
|
||||
4. Add cursor + restart.
|
||||
5. Snapshot as optimization only.
|
||||
|
||||
---
|
||||
|
||||
## 10. Document history
|
||||
|
||||
| Date | Note |
|
||||
|---|---|
|
||||
| 2026-08-03 | Initial design from gm_mining_pool / generations encoding discussion |
|
||||
Reference in New Issue
Block a user