Archival versus pruned
When to use this: you need to decide whether your Generator keeps all historical data or prunes old blocks. A pruned Generator uses less storage but can still serve snapshots for Host bootstrap. An archival Generator keeps everything but grows without bound.
These scenarios use Ethereum Mainnet and Geth. Ethereum is the only officially supported chain today, but any EVM-compatible chain should work by changing chain.name and pointing the RPC URLs at a compatible node. See the chain config for details.
Topology
The pruned Generator writes snapshot files before deleting old blocks. Hosts can download those snapshots for historical bootstrap, then receive live blocks over P2P. The archival Generator never deletes anything. It uses more disk space but does not need snapshots because it can serve any block range directly over P2P.
Pruned Generator config
This config keeps 1000 blocks in the database and produces snapshot files before pruning. The pruner and snapshot values are drawn from the shipped config.yaml, with docs_per_block filled in from the code default in SetDefaults (it is not present in the shipped file):
chain:
name: "Ethereum"
network: "Mainnet"
defradb:
url: "http://localhost:9181"
keyring_secret: ""
embedded: true
p2p:
enabled: true
accept_incoming: false
bootstrap_peers: []
listen_addr: "/ip4/0.0.0.0/tcp/9171"
store:
path: "./.defra"
block_cache_mb: 512
memtable_mb: 64
index_cache_mb: 256
num_compactors: 4
value_log_file_size_mb: 128
geth:
node_url: "${GETH_RPC_URL}"
ws_url: "${GETH_WS_URL}"
api_key: "${GETH_API_KEY}"
api_key_type: "${GETH_API_KEY_TYPE}"
indexer:
start_height: 0
concurrent_blocks: 1
receipt_workers: 8
max_docs_per_txn: 100
blocks_per_minute: 60
health_server_port: 8080
open_browser_on_start: false
start_buffer: 100
logger:
development: true
pruner:
enabled: true
max_blocks: 1000
docs_per_block: 1000
interval_seconds: 30
prune_history: false
snapshot:
enabled: true
dir: "./.defra/snapshots"
blocks_per_file: 100
interval_seconds: 30
Pruned key values
pruner.enabled: true: Turn on automatic pruning. See pruner config.pruner.max_blocks: 1000: Keep the last 1000 blocks in the database. At roughly 12 seconds per Ethereum block, that is about 3.3 hours of history. See pruner config.pruner.docs_per_block: 1000: Average docs per block. The pruner triggers atmax_blockstimesdocs_per_blockdocuments. See pruner config.pruner.interval_seconds: 30: Check for pruning every 30 seconds. See pruner config.pruner.prune_history: false: Do not walk DAG chains to delete historical block versions. Setting this to true is two to three times slower. See pruner config.snapshot.enabled: true: Produce snapshot files before pruning. Hosts can download these for historical bootstrap. See snapshot config.snapshot.blocks_per_file: 100: Bundle 100 blocks per snapshot file. The shippedconfig.yamlsets 100, while the code default inSetDefaultsis 1000. See snapshot config.
Archival Generator config
This config disables both pruning and snapshots. The database grows without bound. Mount this file at /app/config.yaml:
chain:
name: "Ethereum"
network: "Mainnet"
defradb:
url: "http://localhost:9181"
keyring_secret: ""
embedded: true
p2p:
enabled: true
accept_incoming: false
bootstrap_peers: []
listen_addr: "/ip4/0.0.0.0/tcp/9171"
store:
path: "./.defra"
block_cache_mb: 512
memtable_mb: 64
index_cache_mb: 256
num_compactors: 4
value_log_file_size_mb: 128
geth:
node_url: "${GETH_RPC_URL}"
ws_url: "${GETH_WS_URL}"
api_key: "${GETH_API_KEY}"
api_key_type: "${GETH_API_KEY_TYPE}"
indexer:
start_height: 0
concurrent_blocks: 1
receipt_workers: 8
max_docs_per_txn: 100
blocks_per_minute: 60
health_server_port: 8080
open_browser_on_start: false
start_buffer: 100
logger:
development: true
pruner:
enabled: false
snapshot:
enabled: false
Archival key values
pruner.enabled: false: No pruning. The database retains every block, transaction, log, and signature fromstart_heightonward. See pruner config.snapshot.enabled: false: No snapshots. An archival Generator does not need snapshots because it never deletes blocks. Hosts can replicate the full history over P2P. See snapshot config.
Storage comparison
The pruned Generator with max_blocks: 1000 and docs_per_block: 1000 holds roughly 1 million documents at steady state. On Ethereum mainnet, each block produces approximately 150 documents on average (blocks with many transactions produce more, empty blocks produce fewer). The actual document count varies, but 1000 blocks typically uses 2 to 5 GB of Badger storage.
The archival Generator grows by roughly 150 documents per block, or about 650,000 documents per day. At Ethereum's 12-second block time, a year of mainnet history is approximately 2.6 million blocks, producing hundreds of GB of storage. Plan disk capacity accordingly.
Compose file
Either config can be mounted with the same compose file. Change the config file path as needed:
networks:
shinzo-net:
driver: bridge
services:
shinzo-generator:
container_name: shinzo-generator
platform: linux/amd64
image: ghcr.io/shinzonetwork/shinzo-generator-client:standard
user: "1001:1001"
restart: unless-stopped
networks:
- shinzo-net
mem_limit: 16g
mem_reservation: 13g
ports:
- "9171:9171"
volumes:
- ~/shinzo-data/defradb:/app/.defra
- ./config.yaml:/app/config.yaml:ro
environment:
- GETH_RPC_URL=http://localhost:8545
- GETH_WS_URL=ws://localhost:8546
- DEFRADB_KEYRING_SECRET=pingpong
- GOMEMLIMIT=14GiB
logging:
options:
max-size: "50m"
max-file: "3"
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8080/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 60s
Gotchas
pruner.prune_thresholdappears in the shippedconfig.yamlwith value 1. It is deprecated and unused by the pruner. It has been omitted from both configs above. See pruner config.snapshot.blocks_per_filedefaults to 1000 in the code (SetDefaults), but the shippedconfig.yamlsets 100. If you omit the field entirely, you get 1000. If you mount the shipped config, you get 100.- When switching from pruned to archival, old pruned data is gone. The Generator will not re-fetch blocks that were already deleted. You need to restart with
start_heightset to a block you want to begin archival from, or accept that history before the switch is incomplete. - When switching from archival to pruned, the pruner will start deleting old blocks on the next interval. Snapshot files are created before pruning begins, but if the snapshot directory does not have enough disk space, snapshots may fail and blocks will be pruned without a snapshot backup.
- The pruner uses
DefaultCollectionConfigwhich prunesBlock,BlockSignature,AccessListEntry,Log, andTransactioncollections. It does not pruneSnapshotSignaturedocuments, which accumulate over time even on a pruned Generator. - For an archival Generator, consider increasing
block_cache_mbandindex_cache_mbto keep more of the larger database in memory. See the high-throughput tuning scenario for cache sizing guidance.
Need help
- For onboarding and technical support, join the Shinzo Discord.
- To report a documentation bug or request a feature, open an issue in the docs repo.
- For a technical issue with the Generator client, open an issue in the shinzo-generator-client repo.