Skip to content

feat(bitcoin): Recommend Graviton4 (r8g) instance types - #341

Merged
frbrkoala merged 11 commits into
mainfrom
feat/bitcoin-graviton-instances
Oct 2, 2026
Merged

frbrkoala merged 11 commits into
mainfrom
feat/bitcoin-graviton-instances

Conversation

@forrest-not-gump

@forrest-not-gump forrest-not-gump commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Switches the Bitcoin blueprint's recommended and sample instance types from x86 r7i to Graviton4 (r8g), based on side-by-side mainnet tests. It also adds "choose this when" and load-sizing guidance, and corrects several README statements that turned out to be wrong during testing.

Recommendation

r8g.2xlarge (primary) r7g.2xlarge (secondary) r7i.2xlarge (x86)
Hourly price vs r7i (us-east-1 on-demand) -11% -19% —
Sync time, block 400k → tip (two syncs; r7g one) 7.8–8.5 h 8.5 h 8.9–9.1 h
Full-block RPC, single client / under load 7.7/s / 81/s 6.3/s / not tested 7.4/s / 50/s
Tx lookups, single client / under load 2,806/s / 25,000/s 1,812/s / not tested 1,003/s / 17,000/s
  • r8g: won every test. Under concurrent load it served 1.5–1.7× r7i's peak throughput, about 1.8× per dollar. Its 8 physical cores beat r7i.2xlarge's 4 cores with hyperthreading once clients exceed r7i's core count.
  • r7g: the fallback where r8g isn't available, or where the lowest hourly price matters most. It handles all tested workloads. r7i beats it on single-client full-block RPC by about 16%.
  • r7i: only when x86 is needed on the host.
    • Its single-client light-RPC gap comes from vCPUs entering the C6 idle state (190 µs exit latency) between requests. Graviton doesn't expose C-states to the OS.
    • Disabling C6 raised r7i's single-client light-RPC throughput 2.3×. Under load the effect fades, because busy cores stay out of C6.
  • After sync: r8g.xlarge reaches about half of r8g.2xlarge's peak throughput at half the price. It's the most cost-efficient choice for lookup-heavy nodes, and resizing via .env + cdk deploy is in place with no re-sync.

Changes

Samples: mainnet single-node and HA use r8g.2xlarge, and testnet uses r8g.xlarge, all with CPU_TYPE="ARM_64". No node.sh change is needed; it already installs the aarch64 Bitcoin Core build.

README:

  • New "Choosing an instance type" guide: the table above plus "choose this when" guidance.
  • New "Under concurrent load" section:
    • Peak throughput and p99 latency for r8g/r7i at 2xlarge and xlarge.
    • Which workloads are CPU-bound.
    • The network-baseline limit on sustained full-block serving: about 8.4 MB of JSON per block, so roughly 55 blocks/s sustained on r8g.2xlarge.
  • New "After initial sync" section: xlarge sizing guidance, and resizing in place via .env + cdk deploy.
  • Corrections:
    • Mainnet size is ~970 GB and growth ~100 GB/year, measured from the last year of blocks (previously ~650 GB and ~80 GB/year).
    • Measured initial sync time is about 8–10 h (previously 12–48 h).
    • "Upgrading Client Versions": a single-node redeploy with a new CLIENT_CONFIG stops and starts the same instance without re-running setup, so the node keeps the old version. Refs Single-node redeploy with new CLIENT_CONFIG doesn't upgrade the client #340.
    • Changing CPU_TYPE on an existing stack fails on the volume attachment and rolls back.
    • The RPC secret name is <stack-name>/bitcoin_rpc_credentials, not bitcoin_rpc_credentials.
    • The HA cdk destroy stack name was wrong (no -ha suffix).
  • Cleaning Up: warns that cdk destroy deletes the chain data volume, and documents deleting the RPC secret, which cdk destroy leaves behind.
  • Troubleshooting: new entry for a node that crash-loops after an interrupted first boot (re-run node.sh).

Blueprint: defaultDataVolumes is now 1500 GiB, matching the samples.

Deploy prompt: docs/ageai-deploy-prompt.md Step 2.5 now lists bitcoin as a built-in blueprint, so assistants don't demand an external-blueprint security review for it. This is a bug fix.

CHANGELOG: entries under [Unreleased].

Testing

Build: npm run build and npm test pass (22 suites, 473 tests). pre-commit passes on the changed files. npx cdk synth on all three samples produces the Ubuntu 24.04 arm64 AMI with the expected r8g instance type.

Mainnet nodes: every node used:

  • the same blueprint and sample
  • Bitcoin Core v31.1
  • 1.5 TB gp3 at 6,000 IOPS / 400 MB/s
  • us-east-1a

What was measured:

  • Initial sync to tip, two syncs per r8g and r7i and one for r7g, timed per phase from the per-minute c1_block_height metric.
    • r8g was faster than r7i in both syncs: by 14% and 4%.
    • On the multithreaded stretch above assumevalid, r8g was about 30% faster in both.
    • The second r8g node was deployed from this PR's updated sample.
  • Single-client RPC benchmark: two runs per node over one sequential localhost connection, with the same fixed-seed sample on every node (getblock verbosity 2 and getrawtransaction verbose), both cold and warm.
  • Concurrency test:
    • Run from a separate c7g.4xlarge load generator per node in the same AZ, with 1–64 closed-loop clients at 45 s per level.
    • Three workloads: getblock v2, getrawtransaction verbose, and a 90/10 mix.
    • Node CPU was sampled alongside, at 2xlarge and again at xlarge after an in-place cdk deploy resize.
    • 0 errors and 0 HTTP 503s at every level, and the load generator was never CPU-bound.
  • r7i light-RPC profiling: split per request into HTTP/JSON-RPC overhead, txindex lookup and decode, plus host microbenchmarks and a C6-disabled rerun.

Other checks:

  • Resize and redeploy:
    • In-place resizes via cdk deploy worked on Graviton and x86: r7g.2xlarge → xlarge, r8g.2xlarge → xlarge and r7i.2xlarge → xlarge, each with the same instance and volume and no re-sync.
    • A redeploy with changed user data updated the instance in place without re-running setup.
    • A CPU_TYPE redeploy rolled back.
  • Teardown: cdk destroy deleted the data volumes and left the RPC secrets behind.

dbcache 24 GiB vs 4096: not adopted. It was 4.2% faster through block 600k, but the gain was shrinking and the difference was the same size as run-to-run variance. Its hourly UTXO flushes also grew with the cache. The blueprint keeps dbcache=4096.

Caveats:

  • one or two instances per type
  • on-demand pricing in us-east-1 only
  • testnet guidance extrapolated from mainnet
  • r7g not load-tested
  • HA rolling updates not tested
  • the tx-lookup plateau below 100% CPU is inside Bitcoin Core (likely cs_main or the 16-thread RPC pool) and wasn't isolated

The benchmark scripts aren't included in the repo; they're specific to this test and would need generalizing before publishing.

The README and the blueprint's defaultDataVolumes said 1 TB / 1000 GiB
while every mainnet sample provisions 1500 GiB. Update the diagram,
instance and storage tables, cost notes, and package.json to 1.5 TB so
the AI deploy workflow and manual readers see what actually deploys.
Step 2.5 of the AI deploy workflow omitted bitcoin from the built-in
list, so assistants would require an external-blueprint security
review for it. Match docs/ageai-blueprint-security-review.md.
Switch the mainnet (single-node, HA) and testnet samples from r7i to
r8g (CPU_TYPE=ARM_64). In a side-by-side mainnet test with Bitcoin
Core v31.1, r8g.2xlarge beat r7i.2xlarge on sync time (-14%),
full-block RPC (+4%), and transaction lookups (2.8x) at 11% lower
hourly cost. node.sh already installs the aarch64 build, so no code
changes are needed.

Add a "Choosing an instance type" guide (r8g primary, r7g low-cost
secondary, r7i for x86 hosts, including why r7i trails on light RPC:
default C6 idle wakeups), document the expected GuardDuty
BitcoinTool.B finding with a scoped suppression example, and update
the measured chain size and IBD time.
Keep the blueprint docs focused on instance guidance for now; the
GuardDuty finding note will be handled separately.
Reframe r7g as the regional fallback that handles all tested workloads
and state plainly where r7i beats it. Add tested post-sync guidance
(r8g.xlarge matched r8g.2xlarge on single-client RPC at half the
price), note that changing CPU_TYPE on an existing single-node stack
rolls back, correct mainnet size and growth (~970 GB, ~100 GB/year
measured from the last year of blocks), fix the HA destroy stack
name, and warn that cdk destroy deletes the chain data volume.
A redeploy with changed user data (as a new CLIENT_CONFIG produces)
stops and starts the same single-node instance without re-running
node setup, so the node keeps its previous Bitcoin Core version; the
README said the instance is replaced and upgraded. Verified on a
mainnet test stack.

Add a troubleshooting entry for a node that crash-loops after an
interrupted first boot: re-run node.sh from SSM.
Changing INSTANCE_TYPE and running cdk deploy on a single-node stack
stops and starts the same instance with the new type and keeps the
data volume attached; the node resumed at the tip with no re-sync
(verified r7g.2xlarge -> r7g.xlarge on mainnet). Replace the manual
EC2 console resize guidance, which left the stack drifted.
node.sh stores the RPC credentials as <stack-name>/bitcoin_rpc_credentials,
but the README retrieval, troubleshooting, and RPC Authentication text
used bitcoin_rpc_credentials, which does not exist. Because the secret
is created at first boot rather than by CloudFormation, cdk destroy
leaves it behind; document deleting it in Cleaning Up. Both verified
while tearing down the mainnet test stacks.
Reflect the r7g secondary positioning and link the single-node upgrade issue (#340).
@github-actions

Copy link
Copy Markdown

Dependency Review

✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.

Scanned Files

None

@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Scan for commit: 8e7b9c3 | Updated: 2026-10-02 05:44:18 UTC

🔒 Security Scan Results

Scanned files: 7

✅ No security issues found.

forrest-not-gump and others added 2 commits October 1, 2026 08:25
Add an "Under concurrent load" section from a 1-64 client test run
from a separate load generator: r8g.2xlarge served 1.5-1.7x the peak
throughput of r7i.2xlarge on full-block, transaction-lookup, and mixed
workloads, and each xlarge reached about half its 2xlarge. Note the
network-baseline limit on sustained full-block serving.

Replace single-run sync figures with the range from a second sync
(r8g 4-14% faster than r7i), and replace the untested "high-volume"
claim with measured numbers.
Signed-off-by: Nikolay Vlasov <frbrkoala@users.noreply.github.com>

@frbrkoala frbrkoala left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks a lot for the contribution!

@frbrkoala
frbrkoala merged commit 48ec40d into main Oct 2, 2026
7 checks passed
@frbrkoala
frbrkoala deleted the feat/bitcoin-graviton-instances branch October 2, 2026 05:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants