Why WhatsApp and Discord bet on the Erlang VM — and when I wouldn’t
Elena Vasquez
September 18, 2026
WhatsApp famously kept a tiny engineering team in front of a disgusting number of connections. Discord has talked, for years, about Elixir on the BEAM for the parts of the product that are really a lot of little processes talking. Neither company bet on Erlang because the syntax is charming. They bet on a runtime that treats failure as a local event and a socket as something you can afford to own by the million.
I have put production traffic on the BEAM. I have also walked away from it. The walk-away is the part most essays skip, so I will not skip it. Here is why those bets made sense, and the constraints that would make me pick Go, the JVM, or something ruder.
What the VM actually sells
The Erlang VM — the BEAM — is a scheduler plus a process model plus a distribution story. A process is not an OS thread. It is a cheap isolated mailbox with its own heap. You can have hundreds of thousands of them. They crash. A supervisor restarts them. The rest of the node does not take the crash as a suggestion to die.
That model matches a chat system the way a thread pool does not. A user connection is a process. A room can be a process. A guild, a presence shard, a rate limiter — processes. When one user sends a poison message or a slow client falls off the earth, you want that to be their problem. You do not want a shared mutable session table and a prayer.
WhatsApp’s public lore is the extreme version: a small team, FreeBSD, Erlang, a culture of staying close to the VM. Discord’s public lore is more Elixir — same VM, friendlier language, a huge fan-out problem for presence and messaging. Different companies, same shape of load: many live sessions, lots of fan-out, failure that must not become a process-wide event.
I care about this more than I care about “functional programming.” Pattern matching is nice. Hot code upgrade is a party trick I have used exactly once in anger. The product is isolation and preemptive scheduling of tiny work. If your problem is not that, you are buying a costume.
The failures you can survive
Let-it-crash is not sloppiness. It is a refusal to write the ten-thousand-branch defensive style that Java services accumulate when every exception is a policy. You write the happy path. You let the process die. You let the supervisor decide whether to restart, and how many times, and whether the sibling processes should die too — one_for_one versus rest_for_one versus one_for_all. That tree is the design.
I have seen this save a node when a parser hit a bad payload. I have also seen a badly written supervisor tree restart a crash-looping process forever and call it resilience. The VM will not save you from a tree that has no budget. max_restarts exists because optimism is a production bug.
Distribution is the other promise. Nodes connect. Processes can be registered. You can imagine a chat room that lives on one node and is reachable from others. The honest version: distribution is a sharp tool and a source of split-brain stories. I use it inside a cluster I own, with a clear story for netsplits. I do not use it as a replacement for Kafka when the problem is “we need a durable log.” The BEAM is not a durable log. Persistent terms, DETS, Mnesia — I have opinions, and most of them are “use Postgres or a real log unless you know why you are not.”

Elixir, Erlang, and the hiring reality
If I start a BEAM service in 2026 I start it in Elixir unless the existing code is Erlang. Phoenix, LiveView, Broadway, a package ecosystem that a mid-level can enter without learning 1980s telecom vocabulary on day one. Erlang is still the better window into the VM — I want at least one person who can read OTP docs without blinking — but I will not staff a product team on Erlang syntax as a purity test.
Hiring is the first reason I would not bet the company. The BEAM scene is skilled and small. You can hire. You cannot hire like you hire Java. If the company needs a pipeline in three cities, I will not make Elixir the default application language. I will make it the language of the connection fabric, if the fabric is the product, and I will keep the money path on whatever the rest of the staff already debugs at 3am.
WhatsApp and Discord could grow a culture around the VM. They had a problem that made the culture worth it. A ten-person SaaS that invoices once a month does not have that problem. They have a hiring problem they are about to invent.
When I would make the same bet
I would bet on the BEAM when most of these are true:
- The unit of scale is a live session or a mailbox, not a request/response that dies in 40ms.
- Fan-out is the product. Presence, chat, notifications, multiplayer-ish state, a control plane with a lot of agents.
- A single bad actor or a single bad payload must not take the node.
- I can hire or grow two people who will stay long enough to own the supervisor trees.
- I am willing to run observability that understands processes, not only HTTP traces. OpenTelemetry is better than it was. I still want to see mailbox lengths.
A voice/video edge is a maybe. Discord did not put the entire product on Elixir — they have been public about other systems for other shapes. That honesty is the lesson. Use the VM where the process model is the architecture. Do not use it as a brand.
I would also consider Elixir for an internal tool with LiveView if the team already knows it. That is a productivity bet, not a WhatsApp bet. Do not confuse them in a design review.

When I would not
CPU-heavy work. The BEAM is a poor home for tight numerical loops, image codecs, serious crypto you should not rewrite, and ML inference. You can NIFs your way into regret. I put that work in Rust or a native library and I talk to it from the VM, or I do not use the VM.
A CRUD app with a thin websocket afterthought. Phoenix will happily serve that app. So will Rails, Laravel, Spring, FastAPI. I will not pay the hiring tax for a resource-oriented admin because someone liked a conference talk about WhatsApp. If the websocket is the product, stay. If the websocket is a badge, leave.
A shop that wants shared mutable state as a lifestyle. You can write imperative spaghetti in Elixir. People do. If the team’s instinct is a global ETS table for everything, you have lost the isolation story and kept the weirdness. I would rather they write boring Go.
A latency budget that needs a JIT and a profiler culture you already have on the JVM. You can be fast on the BEAM. You can also spend a quarter learning why a process is not getting scheduled the way you thought. If the company already has async-profiler muscle and the problem is a throughput API, I will not relocate that muscle for romance.
When the durable truth is the hard part. Chat history, payments, audit. The VM is great at the live part. The disk part is still disk. I have watched teams invent Mnesia clusters because it felt on-brand. I would take Postgres and a smaller BEAM surface.
A decision I actually ran
We needed presence and fan-out for a collaborative editor — not Discord-scale, but enough that a naive Redis pub/sub plus a Node process per box was shedding clients on deploys. The tempting rewrite was “Elixir, like Discord.” The honest constraint was: two engineers had BEAM time, the rest of the company was JVM, and the document truth already lived in Postgres.
We wrote an Elixir presence cluster for the live fan-out and left the document service in Kotlin. The Elixir nodes were allowed to forget. Postgres was not. Deploys became “restart presence, clients reconnect, documents do not care.” That is the WhatsApp lesson at a smaller budget: isolate the thing that should crash. Do not move the ledger onto the thing that should crash.
I would not have moved the document service. I have seen that rewrite proposed. It dies in the first week of conflict resolution, which is not a mailbox problem.
What I tell leads who want “the Discord stack”
Read what Discord actually said about the parts that are not Elixir. Then look at your connection count, your fan-out, your hiring, and your durable store. If you are chasing a vibe, write the service in the language your on-call already speaks and put timeouts on the sockets.
If you are chasing the VM’s real product — cheap isolation, preemptive scheduling, a supervisor tree you can draw — then bet like WhatsApp did: small surface, people who will live in OTP, metrics that know what a process is, and no shame about the boring database next door.
I would make that bet again for the live fabric. I would not make it for the company default. The companies that look like they bet everything on Erlang usually bet it on a problem that looked like Erlang. That is a different sentence. It is the one I use in design review when someone pastes a WhatsApp slide and calls it an architecture.