N+1, lazy load, and preload: the ORM performance bugs every backend should know by name
James Okonkwo
September 18, 2026
I thought the API was slow. The handler was fine. The serializer was fine. The page issued one query for the invoices and then one query per invoice for the lines, and then one query per line for the tax code. Four hundred statements later, p99 was a confession. The ORM had been helpful. Lazy load had been the default. Nobody had said “preload” because nobody had looked at the log. Every backend that uses an ORM should know these bugs by name. They are not trivia. They are the reason a show page feels like a timeout.
I have hunted this in Hibernate, Active Record, Django ORM, EF Core, SQLAlchemy, and GORM. The names change. The shape does not: N+1, lazy load you did not mean, and a missing preload / includes / select_related / JOIN FETCH.
N+1, said without a whiteboard
You load N parents. Then, for each parent, you load children in a separate query. 1 + N. Sometimes 1 + N + N×M. The code looks clean: invoice.lines.each { |l| l.tax.code }. The log looks like a denial-of-service you wrote yourself.
I turn on the statement log in development. In Rails, verbose_query_logs. In Hibernate, show_sql plus a counter I trust (datasource-proxy, or p6spy, or the Hibernate statistics). In Django, the debug toolbar. If I cannot see the count, I will not see the bug. Dashboards that only show average latency will hide a page that is fine at 3 invoices and dead at 80.
The fix is almost always: load the graph you will touch, in a bounded number of queries. includes(:lines) / prefetch_related / JOIN FETCH / EF Include. Or do not use the graph: write the SQL. I have used sqlc when the query is the product. I have also kept the ORM and preloaded. Both are legal. Pretending the ORM will guess the graph is not.

Lazy load: a feature that is a footgun
Lazy load means “I will query when you first touch the association.” It makes the happy path pretty. It makes a serializer a query engine. It makes a background job that touches a detached entity in Hibernate throw, or worse, silently reopen a session. I have seen both.
I now default new Hibernate entities to lazy on collections — that is normal — and I treat any access outside an explicit fetch plan as a bug. hibernate.default_batch_fetch_size can turn N+1 into a few IN queries. That is a mitigation, not a design. I still want the fetch plan on the path I ship.
In Rails I have used strict_loading on models I care about so a lazy touch raises in development. It is rude. It is how we stopped shipping the show page that only failed in production with a big account. Django’s raise on unfetched relations in some shops is the same rudeness. I want that rudeness.
Lazy load in a GraphQL resolver is a famous cousin. DataLoader exists because N+1 is the default when you think in fields. I use DataLoader or I do not ship the field. I do not “just resolve” a child per parent and hope.
Preload, includes, join fetch — pick the one that matches the cardinality
Join fetch / join: one query, a join, risk of a cartesian blow-up if you join two bags. I join when I need a filter on the child in the same query. I do not join two has-many associations at once and then wonder why I have 10,000 rows for 50 invoices.
Preload / prefetch / IN query: one query for parents, one for children with WHERE parent_id IN (...). Safer for two collections. I use this for the show page that needs lines and events.
Select columns: do not SELECT * a 40-column row if you need three fields for a list. I have seen list endpoints hydrate full entities and then serialize three keys. That is not N+1. It is still a waste. pluck / a DTO query / sqlc list query.
I name the plan in the code. A comment is fine: “preload lines; do not join, two bags.” The next person will thank me. The ORM will not remember why.

How I catch it before the customer does
A test that asserts an upper bound on statements. I have used Testcontainers plus a counter. “This action issues at most 6 queries.” When someone adds a belongs_to in the serializer, the test fails. That test is worth more than a mock of the repository. I said the same in the unit versus integration mix: the risk is the SQL.
Bullet (Rails), New Relic / Datadog APM with N+1 detection, Hibernate statistics in staging, Django debug toolbar in dev. I pick one and I look at it on the fat account fixture, not on the seed with two rows.
I keep a fat fixture. Two invoices will not show N+1 as a latency problem. Two hundred will. I generate them in a spec factory. I do not wait for production.
Serializers, jobs, and the other rooms N+1 hides in
JSON serializers are where I find this most. A presenter that calls invoice.customer.name and invoice.lines.map(&:amount) is a query plan. I preload in the controller or the query object, not in the serializer. If the serializer can load, it will. I have put “no queries here” in a comment and I have enforced it with a counter in the request spec.
Background jobs that iterate a batch and touch associations per row are the night-shift version. I preload the batch. I or I write a join in SQL and I iterate a DTO. Sidekiq making 10,000 queries at 2 a.m. is how you wake up to a database CPU graph that looks like a city skyline.
Admin pages are guilty. ActiveAdmin / Blazer / a Django admin list with a column that is an association. I add includes on the admin queryset. I have forgotten and I have watched a support tool take down the primary. The admin is production. Treat it like production.
Pagination does not save you if you N+1 inside the page. Ten parents still become ten plus children. I still preload the page. I still do not join two bags.
Hibernate hints I still write down
@BatchSize is a seatbelt. @Fetch(FetchMode.SUBSELECT) can be a surprise. Open-session-in-view is a way to hide N+1 until the view. I turn OSIV off on new services. I want the failure in the service layer, not in a JSP I do not have. Entity graphs exist. I use them when the same entity has two fetch shapes. I do not use them as a personality.
Criteria and JPA Specifications that compose dynamically are how you get a query you cannot name. If the page is a report, I write the SQL. I said that in the typed-SQL piece. I mean it here as a performance rule, not as an architecture rule.
What I do not do
I do not disable lazy load globally and then open every association in a panic. I do not cache the world in Redis to hide the queries. I do not add a second ORM. I do not rewrite the app in raw SQL on a Friday because I am angry. I fix the path that pages. I add the counter. I move on.
I do not trust list.size on a lazy bag if I only needed a count. COUNT is a query. size might load the bag. length / count / exists? mean different things in Rails. I read the docs again every year because I forget. Forgetting is how N+1 comes back with a new name.
Every backend should know these names the way they know a unique index. N+1 is the bug. Lazy load is the default that invites it. Preload is the plan. If you cannot say which plan a handler uses, you do not have a plan. You have a hope and a log you have not opened. I open the log. Then I write the plan. Then the API is not slow. It never was. It was 400 queries hiding behind one page.
I still write the names on the PR: N+1, lazy, preload. If I cannot use one of those words, I have not looked. Looking is the skill. The ORM is just the place the skill shows up. I would rather a junior who opens the log than a senior who recites isolation levels and ships a show page that queries in a loop. The loop is the bug. The name is how we stop shipping it twice. Teach the names in onboarding. Put the statement counter in the first backend PR template. The rest of performance work can wait until someone has seen four hundred queries with their own eyes. After that, they will not need a lecture. They will need a preload and a fat fixture. That is the whole curriculum I trust. Open the log first. Name the plan second. Then ship.