You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Companion to:takeover-network-and-capacity-testing.mdAuthor: Denis Angell, Director of Technology
Window: Jul – Dec 2026
Note: This is the time-bound execution plan. The proposal itself is deliberately timeline-free; scheduling lives here.
Author: Denis Angell, Director of Technology, XRP Ledger Foundation
Date: June 23, 2026
Status: Draft for ED / Board review
Re: Two open RFPs — Additional XRPL Network (Test/Staging/Canary) and Capacity Testing Framework
1. Summary
The Foundation has two open RFPs, each budgeted at $100,000, seeking external vendors to (1) stand up an additional mainnet-aligned XRPL network and (2) characterize the XRPL's practical capacity ceiling. I propose the Foundation Technology team take over both milestones and deliver them in-house, reallocating the RFP budget to infrastructure and engineering effort rather than to vendor margin.
This is not a request to build from zero — we have a substantial head start. The Technology team already operates Alphanet (a Foundation-run XRPL network, built and deployed via the sentinel-ai pipeline) and has built much of a performance-testing pipeline of our own: a load generator (xrpld-loadtester), perf-network bring-up tooling, and a sequenced plan modelled on established perf-pipeline methodology. The RFP money should fund finishing and scaling this to the RFP deliverables — not pay a vendor to rebuild capability we are already most of the way to, and would then have to maintain or re-procure.
Owning the milestones in-house gives the Foundation faster delivery, full control of the tooling, reusable assets, lower total cost, and keeps the capability and institutional knowledge inside the Foundation.
2. Why in-house
Driver
Detail
Mandate overlap
These map directly to existing tasks: B2 (staging env for release candidates), B6 (validator performance metrics + baseline report), C8 (5+ network recovery simulations), and B5 (Grafana monitoring for UNL nodes). The RFPs and my objectives are describing the same work.
Substantial head start (already built)
We are well down the path, not starting cold. We operate Alphanet (a Foundation-run XRPL network, built/deployed via the sentinel-ai pipeline), and have a redesigned load generator (xrpld-loadtester: web API, campaigns, genesis fixtures), perf-network bring-up tooling, prefunded-genesis tooling targeting 10–30M-account state, a sequenced perf-pipeline plan, and a set of in-flight node performance-optimization branches. The pieces still in progress — node-metrics export (the DatagramMonitor/XDGM port) and multi-region deploy — are exactly what the RFP would fund finishing.
Existing tooling
xrpld (core node), xrpld-lab (perf-network bring-up + DR-drill fault injection: stall/restart/amendment), xrpld-loadtester (load generation), the logmon trace-log pipeline (tail → parse → DuckDB), the sentinel-ai build/deploy server, the Alphanet deploy flow, and the distribution/Nexus infrastructure. Node-side metrics export (DatagramMonitor/XDGM) is being ported, not yet in the build.
Cost
A vendor award spends most of the $200k on vendor margin and labour. In-house there is no margin: the same $200k funds the build itself — cloud/hosting, public infrastructure, and engineering effort (including a contract engineer, already scoped under C3) — and the resulting assets stay with the Foundation.
Control & reuse
A vendor-built network/framework becomes a dependency to maintain or re-procure. In-house, the framework is re-runnable after every protocol change (an explicit RFP deliverable) and the network is operated by the team that runs the UNL process.
Speed
No RFP evaluation, contracting, or vendor ramp. We can start against the existing tooling immediately.
3. Milestone 1 — Stage / Canary Network
Source: RFP – Additional XRPL Network (Test, Staging, or Canary), $100k build + separate annual maintenance.
Recommended model: a Canary network. Per the RFP the model is deliberately unspecified; a canary is the highest-value choice because it solves the problem none of testnet/devnet do today — exercising release candidates and amendments under mainnet-aligned validator participation before they reach mainnet. This directly serves B1 (UNL repository process: add → staging → monitor → live) and B2.
Our starting point — Alphanet. Alphanet is a Foundation-run XRPL network; builds are deployed to it through the sentinel-ai pipeline (labeled-PR discovery → merge → build → deploy → fund/verify faucet, genesis or non-genesis). Alphanet itself carries no value — it is dev/test infrastructure. It gives us the deploy machinery and operational base; the RFP build is the work of turning that base into a credible, mainnet-aligned canary.
Should the canary carry value? — the central open question. The intuition (and the Kusama precedent) is that a canary is credible because it runs under real economic conditions: real value at stake, run by the people who run mainnet, which surfaces incentive-, fee-, and exploit-class behaviour a valueless network never will. That is a real argument. But it is not obviously right for us, and it deserves a hard look before we commit:
Reimbursement defeats the realism it's meant to create. If the Foundation guarantees users are made whole, users don't actually bear risk — so we get mainnet-lite-with-a-backstop, not authentic economic behaviour. The very thing value was supposed to surface is distorted by the guarantee.
"Capped value" and "real value" are in tension. To bound the Foundation's liability we'd cap value-at-risk; but a low cap isn't realistic economic weight, which undercuts the credibility argument. We cannot simultaneously have meaningful stakes and bounded liability.
Real value on a deliberately unstable network is a soft target. A canary exists to break and is patched constantly. Putting real value on it invites attacks because it's fragile; the reimbursement reserve could be drained repeatedly.
Provisioning value is itself hard and risky. Bridging mainnet value adds custody, security, and legal surface the Foundation may not want to own. A canary-native token only has "value" if a market forms — hard to bootstrap, and it risks becoming a speculative asset with securities exposure. Notably, Kusama's value was emergent (tied to Polkadot's launch and token distribution), not engineered by decree — copying the outcome by fiat may not reproduce the credibility.
Operator reluctance. Even setting the rest aside, mainnet operators may simply not want to run a fast-breaking, value-bearing network that carries reputational and incident-response exposure for them.
Who should run it? — the RFP says the current UNL operators; I think that's a mistake. The RFP's required outcome is ≥15 of the 35 current UNL validators running the canary. Meeting the spirit of that — mainnet-grade validators — is right; having the current mainnet operators themselves run it is not, for several reasons:
Update lag kills the experiment. The canary exists to carry experimental changes — frequent, potentially monthly, updates. Mainnet operators running it as a side concern will be slow to upgrade; if a meaningful fraction lags a version behind, the experimental change never runs under full participation — it stalls, forks, or goes unexercised. The faster we want to move, the worse this gets, and velocity is the whole point.
Experimental code next to production validator keys is a security risk. Running deliberately-unstable, frequently-patched xrpld on mainnet operators' infrastructure puts experimental code adjacent to live mainnet validator keys and operations — an attack surface and operational-contagion risk we should not manufacture.
Burden with no incentive. It adds real exposure for operators with nothing in it for them; participation will be slow and grudging.
It buys little. What we need is a mainnet-grade, representative validator set we can iterate on fast — not the specific current UNL operators. Tying it to them costs velocity and control for marginal added realism.
Recommendation (for Board decision). Two choices here. (1) Value: launch without bridged real value; treat a value layer as a later, justified phase if a specific bug class demands real stakes. (2) Who runs it: use the RFP's own accepted alternative — Foundation-operated mainnet-grade validators plus willing operators with established track records, with the Foundation controlling the update cadence — rather than mandating the current UNL operators. Both preserve what the RFP is actually after (real, mainnet-grade participation; the ability to exercise amendments and release candidates) while keeping the velocity the canary needs and keeping experimental code away from production validator keys. Topology and latency are not engineered into the canary — they are characterized by the capacity/metrics milestone.
Scope (gap from Alphanet → RFP deliverable):
Raise validator participation to mainnet alignment — target ≥15 of the current 35 UNL validators (the RFP default), with the current UNL operators themselves running the canary nodes.
Validator topology mirroring mainnet peering/configuration as practical.
Public-facing infra: RPC endpoint, explorer, faucet, and indexing.
Handover/runbook documentation so the network is operable by the team or transferable.
The value question above resolved with the Board (default recommendation: no real value at launch).
Operating a canary — it moves fast and breaks. Independent of the value question, a canary needs rapid iteration:
Fast build/release turnaround. sentinel-ai already automates build and deploy of xrpld, so we can produce and roll back builds quickly — the velocity a fast-breaking canary needs, and one a vendor engagement would not match.
Coordinated upgrades across operator-run nodes. Because the current UNL operators run their own canary nodes, pushing a fix is a coordination problem, not an automatic one. We will run a defined staged-rollout process with cross-validator config sync (the control whose absence caused the Holesky-Pectra incident) and direct operator comms, so a fix reaches the set quickly.
Validator recruitment: leverage existing UNL operator relationships and the new UNL repository process. Where 15 UNL validators is not immediately achievable, supplement with UNL candidates / operators with established track records (the RFP's accepted alternative).
Budget use: the $100k build envelope funds hosting and the public infra (RPC/explorer/faucet/indexing). Ongoing annual maintenance is quoted separately, as the RFP specifies. If the Board elects a value-carrying canary, the reimbursement reserve is also funded separately, not from the build.
Approach: build on our existing perf-network tooling and xrpld-loadtester to determine the XRPL's practical capacity ceiling under mainnet-like conditions — the two-dimensional sweep the RFP asks for: transactions-per-ledger throughput × ledger-state size. The load generator drives transactions at the network until consensus or processing degrades; the funded work turns this into a documented, re-runnable framework with a published report, and completes the node-side metrics export (the DatagramMonitor/XDGM port) needed to capture where and why it degrades — including the "40M account fall-over" question.
Scope — what the RFP asks, and whether we agree:
Test environment "recreating mainnet conditions" (recent mainnet snapshot, UNL-sized set on mainnet-spec HW, mainnet-like peering).✗ Pushback — we are not doing this. The whole point of the state sweeps is to go past mainnet (10–40M accounts vs ~5.5M today), so a mainnet snapshot can't be the starting state — we hand-generate prefunded genesis state instead. And we don't faithfully recreate mainnet hardware/peering/topology; we run a controlled perf cluster, and characterize the real topology + latency separately (measured on mainnet, reproducible in a lab). The result is method-comparable, not a mainnet replica, and the report must say so plainly.
Throughput sweeps — push tx/ledger until consensus/processing degrades.✓ Agree. This is the core of the engagement.
State sweeps — accounts/trustlines/offers/AMMs/MPTs/escrows well above mainnet, incl. the "40M account fall-over."✓ Agree — and it is precisely why we synthesize state rather than snapshot mainnet (see above).
Transaction mix reflecting mainnet, varyable per type.✓ Agree, with a caveat. We approximate the current mainnet mix; replaying real mainnet traffic deterministically is hard (XRPL has no replay harness), so it's a representative synthetic mix, varyable to isolate per-type cost — not a literal replay.
Measurements.✓ Agree, but the node-side export (DatagramMonitor/XDGM) is being ported and isn't in the build yet, and we measure across layers (state I/O, consensus, relay, fee) rather than assume where the ceiling is.
Re-runnable, documented framework.✓ Agree — and it is the main durable asset, and the core of the in-house argument.
This is the natural extension of C8 (recovery simulations) and B6 (validator perf baseline), and the report feeds C4 (Foundation Protocol Priorities) and B8 (≥1 protocol improvement for security/availability).
5. Combined plan
Stage / Canary Network
Performance Network
Owner
Technology team (Denis) + contract engineer (C3)
Technology team (Denis) + contract engineer (C3)
Already built
Alphanet (Foundation-run), deployed via sentinel-ai
xrpld-loadtester, perf-network tooling, sequenced perf-pipeline plan
Sequence: the capacity framework leads (it is the most built and delivers first results fastest); the canary network proceeds in parallel as the validator set is stood up.
6. Budget
Total $200k of allocated RFP build budget reallocated to: cloud/hosting for both networks, public infra (RPC/explorer/faucet/indexing), and engineering effort (including a contract engineer, scoped under C3). No vendor margin; assets and knowledge stay in the Foundation.
Ongoing annual maintenance is quoted separately from the build budget, as the RFP specifies. If the Board elects a value-carrying canary (§3), a reimbursement reserve is also funded separately — not from the build.
7. Risks & mitigations
User loss from an exploit (only if the Board elects a value-carrying canary) — the value question and its objections are in §3; if value is carried, mitigate by capping value-at-risk and ring-fencing a Board-set reimbursement reserve. The default recommendation (no real value at launch) avoids this risk entirely.
Network breaks / needs urgent patching under load — sentinel-ai automates our build/deploy turnaround so fixes are produced fast; pushing them to operator-run nodes is a coordination problem, mitigated by a defined staged-rollout + cross-validator config-sync process and direct operator comms (config drift across clients caused the Holesky-Pectra incident).
Validator recruitment to ≥15 UNL — mitigate via existing operator relationships and the RFP's accepted alternative (UNL candidates / track-record operators).
Team bandwidth — mitigate via the C3 contract engineer; the heaviest tooling already exists.
Perceived loss of external rigor / independence — mitigate by publishing the capacity report and framework openly (the RFP already expects published results).
8. R&D — open research we need to do
The capacity work depends on research questions that aren't yet answered. These are framed as explicit R&D, run across the three local repos — xrpld (instrumented core node), xrpld-lab (simulation / perf harness), and xrpld-loadtester (load generation / measurement) — and feed Milestone 2's report and the protocol-level recommendations.
Measure mainnet topology & latency, then simulate it in the lab. Map the real inter-validator topology and peer-to-peer latency on mainnet (overlay crawl + a live propagation monitor), then reproduce that topology and latency profile in the lab so sweep results reflect actual network conditions rather than an idealized single-region cluster. This is the foundation that makes the capacity ceilings credible — geo-distributed reality routinely collapses headline numbers by 20×+ versus single-LAN demos.
Get the max-objects dataset from Brett. Brett (ED) holds the max ledger-object figures. Use that as the upper bound / input to the state sweeps.
Correlate object growth with latency, payload, and state I/O. Measure how increasing ledger object counts (accounts, trustlines, offers, AMMs, MPTs, escrows) drives (a) consensus/propagation latency, (b) the per-ledger payload shared between validators (proposals, transaction sets, ledger deltas, state diffs), and (c) per-node state-access (disk) latency. The verified L1 evidence is that for state-heavy chains the dominant scaling bottleneck is state/storage I/O, not CPU — so node-side I/O is a prime suspect for the XRPL ceiling as state grows, alongside the network and latency effects below.
Project against real-world data-transmission trends — bandwidth vs. latency. This is the same account-size-and-latency question pushed forward in time. Two different physics govern it:
Bandwidth grows; it absorbs bigger payloads. Data-center and backbone bandwidth — the layer validators actually run in — compounds at ~25–32%/yr (≈10× per decade; per-link DC Ethernet ~6× per decade: 10G→100G→400G→800G→1.6T). So a heavier per-ledger payload from a larger account/object set becomes proportionally cheaper to ship every year, provided the protocol's data-sharing volume grows no faster than bandwidth.
Latency does not. Propagation is floored by the speed of light (~5 µs/km in fiber; ~70 ms transatlantic RTT, ~150–225 ms US↔Asia) and real inter-region links already run within ~1.3–3× of that floor — there is essentially no headroom, and no amount of money or future hardware lowers it.
Implication for the ceiling. As state grows, payload size is a bandwidth problem (tractable) but consensus round-time is a latency problem (fixed by geography and the number of sequential round-trips). The binding capacity ceiling therefore trends toward latency, not bandwidth. The protocol levers that matter: minimize round-trip count in consensus, keep the latency-critical payload small while letting bulk data ride the bandwidth-rich/latency-tolerant path, and manage validator geographic topology. This is what the R&D must quantify for the XRPL specifically.
9. Methodology — what we're solving, and where it might not work
Below, for each practice the most rigorous L1 programs follow, we state the problem it solves and then — honestly — why it may not hold for the XRPL specifically. These are methodology questions the engagement has to resolve, not settled answers.
Reproducing mainnet conditions ("shadow mainnet").What we're solving: synthetic, batched, no-op load over-reports throughput; starting from real mainnet state and replaying a real transaction mix gives a ceiling that means something.
Why it might not work for us: the whole point of the state sweeps is to go well past mainnet — 10–40M accounts vs ~5.5M today — so beyond mainnet scale there is no real state to copy and we are forced to synthesize/hand-generate it anyway; "don't fake it" can't be taken literally. And the shadow-fork tooling that makes this clean is Ethereum-specific; XRPL has no equivalent deterministic state/tx-replay harness, so adopting the approach is itself part of the build, not a free input.
Where the ceiling actually is — state I/O vs consensus.What we're solving: knowing which layer to instrument first so we measure the real bottleneck, not a convenient one.
Why it might not work for us: the strongest external finding — that state/storage I/O, not CPU, is the dominant ceiling — comes from Ethereum's execution model. XRPL's architecture (SHAMap, NodeStore, the consensus/relay path) is different, so the binding constraint could just as plausibly be consensus round-trips or transaction relay. Treating "I/O is the ceiling" as the answer would bias the instrumentation; we should measure across layers and let the data say, not assume.
Finding the ceiling: worst-case gating and committed goodput.What we're solving: averages with warm caches flatter the result; reporting worst-case ledgers and counting validated (not merely submitted) transactions keeps us honest.
Why it might not work for us: "worst-case" is only meaningful once we define the XRPL degradation signal precisely (round overrun? late validations? publish-latency threshold?) — and that definition is itself unsettled and consequential. Push it too far and the pathological block is so unrepresentative it tells us nothing actionable. The discipline is right; the threshold is a real design decision, not a given.
Real validator participation as "credibility."What we're solving: a network that mainnet validators actually run carries weight; the RFP requires ≥15 of 35 UNL.
Why it might not work for us: the "participation = credibility" lesson is drawn from Ethereum/Polkadot proof-of-stake, where validators have staked economic skin in the game. XRPL's UNL is a trust list, not stake — so the analogy may not transfer, and credibility for the XRPL might rest on operator identity and config fidelity rather than raw validator count. Separately, persuading 15 mainnet operators to run an extra, deliberately-unstable network is a heavy ask with no obvious incentive; if recruitment stalls and we fall back to candidates/track-record operators, the "mainnet-aligned" claim weakens — the very thing the target was meant to guarantee.
Reproducing topology and latency in a lab.What we're solving: an idealized single-region cluster hides the bottlenecks that real geographic spread creates; measuring mainnet topology/latency and replaying it makes the sweep credible.
Why it might not work for us: the emulation has real limits — tc/netem models only a single path and "breaks down" against real multi-level network complexity; Shadow/Kollaps reproduce a topology faithfully but only as faithfully as the matrix we feed them. The honest test is whether emulated propagation percentiles actually match observed mainnet — if they don't, the lab numbers don't transfer, and we'd be back to trusting a geo-distributed live run.
Keeping it from rotting. The recurring failure mode across these programs is rot — testnets left to die, harnesses unmaintained, config drift (the Holesky-Pectra non-finality traced to a testnet-specific config error across clients). This is the argument for in-house ownership with a named owner, and for quoting maintenance as a distinct line item — but only if the Foundation actually funds and staffs that ownership; if it doesn't, in-house rots the same way a vendor deliverable does.
10. RFP coverage & scope discipline
Both RFPs explicitly ask respondents to state what is in scope at the price and what is deferred or excluded. This section maps our response to each RFP's stated requirements and — just as importantly — argues the things we are deliberately not doing, so the $200k buys the highest-value outcome rather than a thin layer across everything.
10.1 Milestone 1 — "Additional XRPL Network" RFP
RFP requirement
Our response
In scope at price
Proposed model + rationale vs alternatives
Canary/staging built on the existing Alphanet; solves pre-mainnet RC/amendment validation that testnet/devnet cannot
✅
≥15 of 35 UNL validator participation (or comparable)
Pushback (§3): RFP wants the current UNL operators running it; we use the accepted alternative instead — Foundation-operated mainnet-grade validators + willing track-record operators, with controlled update cadence
⚠️ alternative
Validator topology + how threshold is met
Topology not engineered to mirror mainnet; measured via the same monitoring (XDGM). Participation met via the alternative model above
✅
Real value bridged / synthetic / both
Open Board decision (§3) — default recommendation is no real value at launch; the case for and against is argued in §3
⚠️ decision
Public infra: RPC, explorer, faucet, indexing, archive, bridge
RPC + explorer + faucet + basic indexing in scope
⚠️ partial (see exclusions)
Deliverable in $100k vs deferred
Participation + core public infra in build; heavier infra deferred
✅
Ongoing annual maintenance
Quoted as a separate line item, not absorbed into build
✅
Reimbursement reserve
Only if a value-carrying canary is elected; then Board-set cap + ring-fenced reserve, funded separately
Pushback (§4): we exceed mainnet scale, so we synthesize prefunded state rather than snapshot mainnet, and run a controlled perf cluster — not a mainnet replica; topology/latency characterized separately
⚠️ method-comparable
Throughput sweeps
xrpld-loadtester "blast mode" to the degradation point
✅
State sweeps (accounts, trustlines, offers, AMMs, MPTs, escrows)
Prefunded-genesis to 10–30M+, incl. the 40M fall-over scenario
✅
Transaction mix reflecting mainnet, varyable
Mainnet-weighted mix, varyable to isolate per-type cost
✅
Measurements (provider proposes)
DatagramMonitor/XDGM export (port in progress): across layers — state I/O, consensus latency, peer/relay volume, per-type latency, fee escalation, CPU/mem — measured, not assumed
✅
Reports with rendered graphics + downloadable raw data
MD/JSON reports + graphics + raw data
✅
Final report: ceilings, failure modes, protocol recommendations
Across consensus, peering/relay, tx handling, fee engine
✅
Re-runnable, documented framework
Scripts + handover docs preserved for re-runs after protocol changes
✅
10.3 What we are deliberately NOT doing — and why
These are conscious scope choices, not gaps. Each protects the budget or reduces risk.
We are not putting real value on the canary at launch (default recommendation). §3 argues why a value-carrying canary is not obviously right for us — reimbursement undercuts the realism it's meant to create, capped value isn't real value, and real value on an unstable network is a soft target. If the Board elects value, it is capped and reserve-backed; the default is to defer value to a later, justified phase.
We are not building a net-new network — we are extending Alphanet. A from-scratch build would duplicate cost and time for a capability we already operate. The funded work is the gap to mainnet-aligned participation and public infra.
No archive nodes or heavy indexing/bridge components in the initial build. Core RPC, explorer, faucet, and basic indexing ship first; archive and deeper indexing are deferred into maintenance or a later phase. Rationale: the participation outcome and developer-usable basics deliver almost all the value within $100k; archival infrastructure is disproportionately expensive to stand up and run.
We are not chasing all 35 UNL validators. ≥15 is the RFP's own threshold and is sufficient for credibility; pursuing the full set adds coordination cost without changing the outcome.
No standing live dashboard in the capacity build. The RFP asks for reports with rendered graphics and downloadable raw data — which we deliver — but a continuously-hosted dashboard is a maintenance product, not a measurement deliverable. Budget goes to measurement fidelity; a dashboard can follow if the Foundation wants it.
We are not creating all large-scale state through real on-chain transactions. Real account creation is only practical to moderate scale; beyond that we use hand-generated genesis preload to reach 10–30M+ objects efficiently. This is a method choice the report will document.
Stated caveat (not an exclusion): geo-distributed, not single-region. RippleX's reference runs were single-region/LAN, so they were CPU/consensus-bound. We run multi-region for realism, which lowers the throughput knee and raises latency. That is the correct result, but the report will state it explicitly so numbers are not naively compared to single-region figures.
11. Decision requested
Approval to withdraw both RFPs from external award and assign the two milestones to the Technology team, reallocating the combined $200k to infrastructure, public infrastructure, and engineering effort. The one decision this proposal does not make for the Board is whether the canary carries real value (§3) — that is put forward as an explicit choice.