Skip to main content
Architecting Software in 2026 · 1 of 6 partsEngineering11 min read

When to Use Microservices in 2026: Scaling Alone Is No Longer the Reason

On the day the Dow posted a then-record gain, Robinhood's customers could not trade from the open to the close. Splitting out the hot part so it scales alone was the old answer; a serverless platform now does that per function. I open this series with Bullpen, a paper-trading league on live market data, and argue when to use microservices in 2026: only for a reason I can name, with a check, a measurement or a bill behind it.

← All Posts
2/4

Robinhood's app went dark at the market open on March 2, 2020, and stayed dark through the close. That day the Dow gained 1,294 points, then a record, and the S&P 500 added $1.1 trillion. Every Robinhood customer, 12.5 million accounts by FINRA's count, could only watch. The lesson most engineers drew from days like that was microservices: split the system so the hot part can scale on its own.

In 2026 that lesson no longer holds on its own: a serverless platform already scales every function with its traffic. When to use microservices now comes down to a different question: what does a split buy, and can I show the evidence before I build it?

This is the first part of a series in which I build one real distributed system, Bullpen, and make every split earn its place with a check, a measurement or a bill. It records my take as of September 2026. The platforms under it change monthly, and parts of it will age. I am starting it the month Sam Newman's Building Resilient Distributed Systems reached early release; its subject, resilience, is what the second reason below asks Bullpen to prove.

Why scaling stopped being enough to split

Robinhood's own account of that morning is worth reading slowly. In their update the next day, the founders wrote that record volume, volatile markets and record sign-ups stressed their infrastructure. The stress set off a thundering herd, and the herd took down their DNS system. FINRA's 2021 settlement, a $57 million fine plus about $12.6 million in restitution, names that outage as the most serious of the period.

My reading is that this was not the scaling failure the microservices pitch had in mind. A dependency that every path shared gave way under a crowd, and everything behind it fell at once. More services calling the same DNS would have failed the same way. A herd is what a broken caller contract looks like at full size, which I wrote about in Backpressure Is a Contract Every Caller Must Honor.

Software Architecture: The Hard Parts names the unit that shares a fate: the architecture quantum, an independently deployable unit with high functional cohesion, held together by static dependencies and synchronous calls. By my reading, services that all wait on one dependency behave as one quantum at runtime, whatever the repository count says.

The scaling half of the old argument moved too. Vercel has enabled Fluid compute by default for new projects since April 2025, and one function instance can serve more than one request at a time. Vercel also bundles Next.js routes into as few functions as it can, and each function adds instances as its traffic grows.

A route that needs different limits gets its own bundle through a functions entry in vercel.json, still inside one deployable. That is the "scale the hot part on its own" split, done as configuration and billed per use. I argued in my notes on backends becoming software that sleeps that the billing model is the tell. Bullpen is where I find out whether that holds under a real workload.

Bullpen, a paper-trading league on live market data

Bullpen is a league where players trade play money on real U.S. stocks and crypto at live prices, and a leaderboard ranks them as the market moves. It runs on free tiers, and the APIs under it fail the way production systems do: partial fills, rate limits, dropped messages.

Bullpen's stock prices come from Alpha Vantage, starting in Part 3, and its crypto prices from CoinGecko. Every free market-data plan rations calls, and that ration is a limit the whole system shares. Alpaca's free market-data plan is a typical example: real-time prices from the IEX exchange only, at most 30 streamed symbols, and 200 API calls a minute. Alpaca's paper trading also sets a useful bar for order realism: it fills orders partially, for a random size, 10% of the time. For crypto, my first pick was Coinbase's public ticker, which needs no key. Coinbase's market data terms forbid showing its prices to anyone outside my own organization, so the terms, not the API, decided the provider. CoinGecko's free Demo plan allows public display as long as every price carries a "Powered by CoinGecko" credit.

Part 1 ships the smallest working version: a landing app, a trading app with one live symbol (BTC-USD from CoinGecko, refreshed every 5 minutes) and its decision records. Both apps live in one Turborepo monorepo. That was my call for this project, and Part 2 argues it. The two apps are not a split in the sense this post cares about: they share packages, hold no data and never call each other. The splits I have to justify are the ones that put a network call between two parts of one order.

One order, four needs, two shared limits

The case for splitting shows up when I follow a single buy through the system.

  • The price feed has to be fresh, and it draws from one upstream budget that every part of Bullpen shares.
  • The order has to reserve buying power before it executes. It may fill in part, and it may race a cancel.
  • The ledger has to keep every balance correct. Buying power can never go negative, even when a fill lands late.
  • The leaderboard has to answer every player at once, hardest at 9:30, and it can tolerate prices that are seconds old.

Those are four different architecture characteristics in Ford and Richards' sense: freshness, correctness under races, strict consistency and read fan-out. Inside one deployable they share everything: the same function instances, the same database connections, the same market-data budget. The diagram below follows the buy from left to right; watch the two shared limits underneath, because that is where the needs collide.

Neither limit bites with one player. Both bite at the open, when every player arrives in the same minute.

The monolith that would carry a classroom league

For a class of 30 students, one Next.js app and one Postgres database would carry Bullpen without trouble. The traffic is small, the market data fits the budget, and nobody notices a two-second delay. For a classroom, I would build that version and stop.

It breaks in three places once the league grows past one room, and the free tiers put a number on each:

Where it breaksFree-tier limitWhat Bullpen needs
Price stream300 s per function on Vercel HobbyA 6.5-hour session: 78 function lifetimes back to back
Market data200 calls a minute on a free plan like Alpaca'sOne budget for all instances: 7 instances polling every 2 s spend it
Database97 usable connections on Neon's smallest computeLedger writes and leaderboard reads from every warm instance at 9:30

The sources are Vercel's function duration limits, Alpaca's plan page above, and Neon's connection pooling docs, which list 104 connections for the smallest compute with 7 reserved. Neon's pooler accepts up to 10,000 client connections, but it still funnels them into that small pool.

Each row points at a boundary. The stream wants something that outlives a function. The market-data budget wants exactly one owner that fans prices out to everyone else. The connection pool wants the ledger's writes protected from the leaderboard's reads. Adding capacity fixes none of them; each fix stops parts with different needs from sharing one fate.

Barry O'Reilly's residuality theory starts by listing the stressors that could hit a system and asks what is left of it after each one. Four are visible on day one:

  • The feed goes down.
  • A sequence gap drops a fill message.
  • The database is cold at 9:30, because Neon's free plan scales it to zero after 5 idle minutes and I cannot turn that off.
  • The market-data budget runs out mid-session.

For each one I care about what a player sees: stale prices labeled as stale, orders paused rather than lost, and no balance that is wrong. What is left after each stressor is the evidence the second reason below asks for.

When to use microservices: the five reasons a split has to earn

Sam Newman put it plainly in Monolith to Microservices: microservices are not the goal. A split is worth it when the current architecture cannot reach a goal I can name. I wrote Bullpen's goals down before building anything, in ADR-0002 of the Bullpen repository. Every future split has to cite one of five reasons and bring a check, a measurement or a bill as evidence:

  1. Data contention. Write paths that should not share a lock or a database. In Bullpen, reserving buying power and settling fills against leaderboard reads. Part 3 brings the evidence.
  2. Failure isolation. One part's outage must not take another down. In Bullpen, a dead price feed must not stop order entry. The bad-morning test in Part 5 is the evidence, and isolation has a price, which I looked at in Cell-Based Architecture Isn't Free.
  3. Independent change. A release cadence that really differs. I expect this one to be the weakest in Bullpen, because one person ships everything. Part 6 counts the deploys. Bullpen also cannot test the reason that matters most in many organizations: teams that need to own and ship their part without waiting on each other. I am the only engineer, so team autonomy is outside this experiment; Part 6 sketches where the team boundaries would sit if Bullpen had them.
  4. Runtime fit. A workload the default runtime serves badly. In Bullpen, a stream that outlives a 300-second function. This is the reason behind the Rust ingestion service later in the series, and it still has to be shown, not assumed.
  5. A shared budget. ADR-0002 frames this as cost split across teams. With one person on the project, it becomes a limit that every instance draws from and only one owner can manage. In Bullpen, that is the market-data call budget, like the 200 calls a minute above.

Each reason justifies a boundary, not a network call. A module, a separate function, a worker behind a queue and a separate service are all boundaries, and I pick the cheapest one that satisfies the reason. Only a boundary that needs its own deploy, its own data and its own failure domain earns a microservice. Runtime fit almost always does, and independent change does by definition. A shared budget almost never does: one module that owns the limiter is enough.

Scaling on its own is missing from that list on purpose. It still splits a system when a workload needs a different runtime, such as a stream that outlives a function, and then it counts as runtime fit. If I cannot name one of the five, the split does not happen. If the evidence never shows up, the split gets merged back.

One more pull to resist: Neon gives every project its own 0.5 GB and 100 compute-hours a month, so a database per service multiplies my free allowance instead of my bill. A cheap pattern is not a reason. Part 3 is about that trap.

The first bill and the free-tier ceiling

Cost is easy to forget from behind an IDE, so Bullpen puts it in the architecture from the first week. On Vercel's Hobby plan, a month includes 4 hours of active CPU, 360 GB-hours of provisioned memory and 1 million function invocations. Cron jobs run at most once a day and fire anywhere inside their hour.

The rule that shapes the design is what happens past those numbers. In most cases, a feature that goes over waits until 30 days have passed before it works again. For a public demo, a paused feature is an outage with a 30-day recovery time. Vercel Queues, which Bullpen will lean on, has been in public beta since February 2026; if its limits change when it leaves beta, this bill changes with it.

So the Bullpen repository keeps a ledger of its own: every free allowance, with a link to the page that states it, in cost/allowances.json, and one usage file per week in cost/usage/. The landing page shows the same ledger in its cost section. Week 40's file stays marked pending until the week closes on October 4, and then gets the real numbers.

The first five days already give a baseline. From September 25 to 30, both apps together used 431 function invocations, 39 seconds of active CPU and 0.04 GB-hours of provisioned memory. None of the three reached 0.3% of its monthly allowance, and the bill was $0.00. That is the easy part: one symbol, one cached call every 5 minutes, and nobody trading yet.

What would change my mind

The question this part leaves open is the one Bullpen forces: does a live trading league on free tiers need to be distributed at all? I do not know yet, and two results would settle it.

  • If the market-open load test in Part 5 shows one deployable serving 1,000 players inside the free tier, with prices fresh enough, I merge the splits back and say so.
  • If the bad-morning test shows a dead feed stopping order entry, failure isolation is proven, and I split sooner than planned.

Until Part 5 runs both tests, the five reasons are a hypothesis, and no split ships in Bullpen without one of them.

Read next

Still here? You might enjoy this.

Nothing close enough — try a different angle?

Architecting Software in 2026 · 1 of 6

Part 2 is in progress.

New parts land on Mondays, 9am Pacific — leave an address and I'll send each one the day it ships. Nothing else.

Was this helpful?

Leave a rating or a quick note — it helps me improve.