Skip to main content

Storage-Layer Reed-Solomon Erasure Coding (10+4)

Membuss integrates protocol-level Reed-Solomon erasure coding (klauspost/reedsolomon) directly into ingestion, block storage, peer placement, transparent retrieval, and background repair pipelines.


1. Adaptive Galois Field Matrix Sharding (GF(2^8))​

Membuss dynamically optimizes the data (K) and parity (M) shard configuration based on content size to maximize fault tolerance while minimizing storage overhead:

Payload SizeData Shards (K)Parity Shards (M)OverheadFault Tolerance
< 64 KB21+50%1 missing shard
< 1 MB42+50%2 missing shards
< 10 MB83+37.5%3 missing shards
≥ 10 MB (Default)104+40%4 missing shards
Original Block Payload (256 KiB)
├── Data Shards (D1 ... D10) : 10 × 25.6 KiB
└── Parity Shards (P1 ... P4) : 4 × 25.6 KiB

2. Ingestion & Manifest Persistence​

During ingestion (AddWithProgress / AddDirectory):

  1. Shard Encoding: Each leaf block is encoded using erasure.NewEncoder(erasure.AdaptiveConfig(size)).
  2. Manifest Creation: An ErasureManifest protobuf structure is constructed linking OriginalMid, DataShards, ParityShards, and ShardMids.
  3. Storage & Announcement: All 14 shard blocks are stored in the Blockstore under their respective ShardMID hashes and announced across the P2P DHT network.

3. Transparent Resilient Retrieval & Reconstruction​

When a client requests a MID via fetchingBlockstore:

  1. Direct Fetch: Attempts to retrieve the primary block directly.
  2. Erasure Fallback: If the primary block is missing or storing peers go offline:
    • Reads the ErasureManifest from store metadata; if absent, fetches it from peers over the manifest RPC.
    • Fetches available shards from connected swarm peers and placement holders (discovered via shard-set records on the DHT).
    • As soon as at least K = DataShards valid shards arrive, executes encoder.Decode(shards, manifest).
    • Verifies reconstructed bytes match the expected BLAKE3 hash.
    • Restores the block locally and streams it seamlessly to the caller.

4. Background Shard Repair Worker​

The repair worker (core/erasure/repair.go) audits sealed MIDs:

  1. Health Audit: Checks presence of all N = DataShards + ParityShards shards for manifest-bearing blocks.
  2. Degraded Detection: Identifies MIDs with K ≤ present < N shards available.
  3. Shard Reconstruction: Reconstructs missing shards via matrix inversion and writes them back locally.
  4. Graceful Skip: MIDs whose manifests carry no shard list (e.g. root summaries) are skipped without error.

See Shard Placement for how shards are distributed to peers in the first place.