Student: Rodrigo Yáñez Pilgrim
Organization: Wagtail (Google Summer of Code 2026)
Mentors: Thibaud Colas, Coen van der Kamp
Project: Scaling translation operations in wagtail-localize
Project page: https://summerofcode.withgoogle.com/programs/2026/projects/wmfuUxir
The main goal of the project is to identify and optimize processes within wagtail-localize that have scalability problems, that is, processes whose cost grows with the size of the site. As a starting point, a preliminary investigation was carried out that identified several areas with potential for optimization.
The areas identified initially were the following:
- Building the page index
- The scalability of translation workflows, including processing page subtrees, saving translations, and moving the heaviest operations outside the HTTP request cycle
- Detecting content changes before updating a translation
During the first two weeks of the project, the proposed flows were diagnosed. The result led us to rethink the project's scope, since in some cases the optimization was possible, but the benefit did not make up for the cost in code readability and maintainability.
A new investigation was carried out, which consisted of building wl-benchmark, a Wagtail project populated with enough volume of pages, articles, snippets, StreamFields and related objects, in multiple languages, to reproduce how a real site behaves at scale. This project was initially created locally as a testing lab. On this site, the extension's main flows were run, such as translating pages, subtrees and snippets, editing sources, syncing, and re-syncing the page tree across locales, measuring how many queries each operation made and how that number grew as the content size increased.
The measurement confirmed the page index problem that was already suspected: the cost grew with the size of the page tree, and a prototype showed it could be reduced to a fixed number of queries, regardless of how many pages there were. That area stayed unchanged in scope.
The other two areas did not have the same result. Content change detection was discarded. The method that was going to be optimized is barely used in the package's real flows — it mostly shows up in tests. The problem did not exist in practice.
The scalability of translation workflows had a mixed result. Grouping the save with bulk_create() was viable, but saving the translation for one locale already made few queries. It complicated the code without a clear benefit, so it was discarded. Moving save_target() outside the request cycle remains unresolved. It is not a query-count problem, it is a response-time problem, and that requires a different kind of measurement.
Looking closely at this area — the processing of subtrees and the full create_translations() flow — turned up a finding that was not on the original list: refresh_segments(), where every text segment extracted from a page fires several chained queries, with a cost that grows fast on pages with many segments. Alongside it appeared a smaller pattern, repeated in two different places: when processing related objects inside create_translations(), and when loading related segments inside save_target().
The pattern confirms what had already been laid out in the scope redefinition: the page index and the related-object fan-out are resolved almost completely (they end up at a fixed number of queries, regardless of size); refresh_segments improves notably but does not reach a fixed number; and save_target and get_or_copy_target barely improve, confirming that they were not the right targets.
Raw output from wl-benchmark: each flow against its own database, freshly seeded with standard content (seed_realistic_localize_universe, default values), with nothing left over from a previous run.
Page index — synctree PageIndex scaling
synctree PageIndex scaling
10 articles 50 articles 100 articles 250 articles
--------------------------------------------------------------------------------------
indexed pages 53 105 170 363
aliases 2 8 15 36
baseline
total queries 260 520 845 1810
from_database queries 260 520 845 1810
entry calls 53 105 170 363
duration 63.6ms 129.6ms 212.4ms 543.3ms
prototype
total queries 2 2 2 2
from_database queries 2 2 2 2
entry calls 53 105 170 363
duration 6.3ms 10.9ms 18.0ms 39.6ms
reduction (total queries)
99.2% 99.6% 99.8% 99.9%
Wrote profiler events: .tmp/benchmarks/synctree-index-scaling.jsonl
refresh_segments — refresh_segments scaling
refresh_segments scaling
1 groups 2 groups 5 groups 10 groups 20 groups
------------------------------------------------------------------------------------------
segments 27 43 91 171 331
baseline cold
cold queries 327 464 944 1716 3281
cold duration 59.4ms 82.5ms 168.8ms 294.0ms 550.2ms
prototype cold
cold queries 100 130 220 370 670
cold duration 34.1ms 46.6ms 87.4ms 149.0ms 290.6ms
baseline refresh
refresh queries 218 364 802 1532 2992
segment queries 187 333 771 1501 2961
max segment queries 10 10 10 10 10
refresh duration 44.8ms 72.7ms 149.5ms 275.7ms 549.1ms
prototype refresh
refresh queries 73 103 193 343 643
segment queries 32 62 152 302 602
max segment queries 2 2 2 2 2
refresh duration 30.5ms 48.1ms 142.3ms 150.0ms 268.9ms
reduction (refresh queries)
66.5% 71.7% 75.9% 77.6% 78.5%
Wrote profiler events: .tmp/benchmarks/refresh-segments-scaling.jsonl
save_target — save_target scaling
save_target scaling
1 groups 5 groups 10 groups 20 groups 50 groups
------------------------------------------------------------------------------------------
baseline (cold)
cold queries 211 221 231 251 311
cold duration 78.0ms 89.1ms 100.5ms 128.5ms 220.7ms
prototype (cold)
cold queries 209 209 209 209 209
cold duration 78.9ms 91.0ms 96.9ms 141.5ms 188.7ms
baseline (warm)
warm queries 159 167 177 197 257
warm duration 52.5ms 58.3ms 77.7ms 103.8ms 175.7ms
prototype (warm)
warm queries 155 155 155 155 155
warm duration 57.7ms 126.1ms 70.1ms 101.1ms 152.7ms
baseline (updated)
updated queries 162 170 180 200 260
updated duration 123.4ms 68.7ms 73.8ms 122.5ms 214.0ms
prototype (updated)
updated queries 158 158 158 158 158
updated duration 55.0ms 71.6ms 69.5ms 100.9ms 150.4ms
reduction (cold queries)
0.9% 5.4% 9.5% 16.7% 32.8%
Wrote profiler events: .tmp/benchmarks/save-target-scaling.jsonl
Related-object fan-out — related object fan-out scaling
related object fan-out scaling
1 groups 5 groups 10 groups 20 groups 50 groups
------------------------------------------------------------------------------------------
related segments 3 7 12 22 52
baseline
total queries 128 30 45 75 165
related queries 126 28 43 73 163
related calls 3 3 3 3 3
create calls 3 3 3 3 3
duration 28.1ms 7.1ms 10.0ms 14.8ms 32.2ms
prototype
total queries 11 11 11 11 11
related queries 9 9 9 9 9
related calls 3 3 3 3 3
create calls 3 3 3 3 3
duration 3.7ms 3.6ms 3.8ms 4.4ms 4.4ms
reduction (total queries)
91.4% 63.3% 75.6% 85.3% 93.3%
Wrote profiler events: .tmp/benchmarks/related-object-fanout-scaling.jsonl
get_or_copy_target — get_or_copy_target scaling
get_or_copy_target scaling
parents-ready
1 groups 10 groups 50 groups
------------------------------------------------------------------
baseline
total queries 211 229 309
copy queries 66 66 66
duration 68.5ms 93.3ms 219.0ms
prototype
total queries 207 207 207
copy queries 66 66 66
duration 71.0ms 88.7ms 232.5ms
reduction (total queries)
1.9% 9.6% 33.0%
parents-missing
1 groups 10 groups 50 groups
------------------------------------------------------------------
baseline
total queries 294 312 392
copy queries 149 149 149
duration 98.8ms 117.4ms 245.6ms
prototype
total queries 290 290 290
copy queries 149 149 149
duration 88.4ms 112.7ms 218.0ms
reduction (total queries)
1.4% 7.1% 26.0%
Wrote profiler events: .tmp/benchmarks/get-or-copy-target-scaling.jsonl
After seeing the results from wl-benchmark, it was proposed to bring this tool into the project on a permanent basis. Several measurement tools were suggested (query-doctor, Silk, django-eagle, django-o11y, nplusone, the pytest family). After analyzing them individually, django-query-doctor was identified as the most useful one for our case.
The harness was designed as a combination of custom tooling together with django-query-doctor. It includes a selection of wagtail-localize flows that showed signs of potential optimizations. In a first phase, it included 8 flows covering 15 executions. While trying to implement the page index case, the need to obtain query attribution came up, so an additional script was created for that.
The harness runs the package's real translation operations — submitting a page for translation, translating a snippet, syncing the page tree, refreshing segments, among others — in an isolated process, against a clean database created for each repetition, without reusing state between runs. It measures query count and time at different content sizes. Each run writes a JSON report with the commit, branch, and versions used, so any number can be reproduced. Queries are the primary result; time is secondary, since it depends on the machine it runs on.
This optimization was initially proposed as a standalone PR (#930), designed based on the measurements taken with wl-benchmark. After the harness was implemented, it was decided to open a new PR built around the integrated tool, so the optimization would have a clear history that would make it possible to understand and reproduce the measurements in the future.
PR #953 runs the harness's three main scripts to get the diagnosis. PageIndex.from_database() cost 5 queries for every indexed page: the parent, the content type, the locale, and one query for each list of locales, the existing ones and the aliased ones. The cost grew linearly with the size of the tree — 119 queries for 24 indexed pages, 319 for 64.
The attribution script pointed to the 5 queries responsible for the growth, each at its line in synctree.py. The fix builds three mappings in a single pass and preloads the two missing relations. Result: from 119 to 2 queries at the small size, from 319 to 2 at the large size. Constant, regardless of the size of the tree.
When compared against query-doctor, the tool only detected 2 of the 5 real sources. The other 3 vary the translation_key on every call, so the query text is never identical and its duplicate detection does not group them. This is the same limitation that had already been anticipated when evaluating the tools.
The PR adds a regression test that checks the query count does not grow as the tree expands. The full suite, 608 tests, passes. One trade-off remains unmeasured: the new version keeps four columns of every page in the tree in memory while it builds the index — the net effect on memory was not measured, and remains as follow-up work.
| Before | After | |
|---|---|---|
| Queries at 24 indexed pages | 119 | 2 |
| Queries at 64 indexed pages | 319 | 2 |
| Growth per page | 5.00 | None |
| Median time at 64 pages | 63.3 ms | 2.8 ms |
Raw output from the three commands (before the fix)
$ python benchmarks/run.py core_page_index --repeat 5
core_page_index [small] 119 queries 24.2 ms median (min 23.6, max 25.3, 5 repeats) 24 indexed_pages (expected 24)
core_page_index [large] 319 queries 63.3 ms median (min 62.0, max 66.4, 5 repeats) 64 indexed_pages (expected 64)$ python benchmarks/run_attribution.py core_page_index
core_page_index: 24 -> 64 indexed_pages
small 119 queries = 119 attributed + 0 outside execute_wrapper
large 319 queries = 319 attributed + 0 outside execute_wrapper
Growth drivers (5 of 5 growing groups)
+1.00/indexed_pages 22 -> 62 synctree.py:72 in from_page_instance (fixed -2)
SELECT "wagtailcore_page"."id", "wagtailcore_page"."path", ...
+1.00/indexed_pages 24 -> 64 synctree.py:75 in from_page_instance
SELECT "django_content_type"."id", "django_content_type"."app_label", ...
+1.00/indexed_pages 24 -> 64 synctree.py:77 in from_page_instance
SELECT "wagtailcore_locale"."id", "wagtailcore_locale"."language_code" ...
+1.00/indexed_pages 24 -> 64 synctree.py:79 in from_page_instance
SELECT "wagtailcore_page"."locale_id" AS "locale" ... WHERE alias_of_id IS NULL
+1.00/indexed_pages 24 -> 64 synctree.py:85 in from_page_instance
SELECT "wagtailcore_page"."locale_id" AS "locale" ... WHERE alias_of_id IS NOT NULL
Largest fixed cost (1 of 1 constant groups)
1 fixed 1 -> 1 synctree.py:152 in from_database
SELECT "wagtailcore_page"."id", "wagtailcore_page"."path", ...$ python benchmarks/run_query_doctor.py core_page_index --size small
core_page_index [small] 24 indexed_pages
119 queries, 2.5 ms of database time under diagnosis
25 prescription(s)
CRITICAL n_plus_one (24 queries)
N+1 detected: 24 queries for table "django_content_type" (field: content_type)
fix: Add .select_related('content_type') to your queryset
at src/wagtail_localize/synctree.py:75 in from_page_instance
CRITICAL n_plus_one (24 queries)
N+1 detected: 24 queries for table "wagtailcore_locale" (field: locale)
fix: Add .select_related('locale') to your queryset
at src/wagtail_localize/synctree.py:77 in from_page_instanceThe remaining 23 prescriptions are 22 duplicate_query warnings of two to four queries each, and one informational note about a missing index on depth in Wagtail's Page model.
Raw output from the three commands (after the fix)
$ python benchmarks/run.py core_page_index --repeat 5
core_page_index [small] 2 queries 1.6 ms median (min 1.5, max 1.7, 5 repeats) 24 indexed_pages (expected 24)
core_page_index [large] 2 queries 2.8 ms median (min 2.6, max 2.8, 5 repeats) 64 indexed_pages (expected 64)$ python benchmarks/run_attribution.py core_page_index
core_page_index: 24 -> 64 indexed_pages
small 2 queries = 2 attributed + 0 outside execute_wrapper
large 2 queries = 2 attributed + 0 outside execute_wrapper
Growth drivers (0 of 0 growing groups)
none
Largest fixed cost (2 of 2 constant groups)
1 fixed 1 -> 1 synctree.py:158 in from_database
1 fixed 1 -> 1 synctree.py:168 in from_database$ python benchmarks/run_query_doctor.py core_page_index --size small
core_page_index [small] 24 indexed_pages
2 queries, 0.2 ms of database time under diagnosis
1 prescription(s)The one remaining prescription is the missing index on depth, which is present before the change as well and belongs to Wagtail's own Page model.
Other contributions came up during the project. They do relate to the original proposal, but they belong to a broader view of the package.
The first was a demo site, as a reference implementation of wagtail-localize. It was proposed in issue #928 and opened as PR #929, the first of the program. Its review also raised whether the project should move to a src layout, so the PR was split in two: #941 for the layout and #942 for the demo.
When that first PR was published, a test failed during CI. It was investigated and found to be related to issue #922, open since before the program and unrelated to the demo: a bug in the translation of snippets with three or more levels of nesting. The fix was merged in PR #934. The same cause affects copy_for_translation() in Wagtail core, so it was reported there in issue #14425, which is still open.
The layout change was merged before the rest of the work. The harness and the page index optimization were already developed on top of the src layout and depend on the new paths, so that change had to land first.
The benchmark harness, PR #946, was merged on August 20, 2026.
The page index optimization went through two PRs. The first, #930, was designed with the wl-benchmark measurements, before the integrated harness existed. It was closed without merging on August 20, the same day the harness merged. It was replaced by #953, which redoes the same optimization on top of the now-finished harness. #953 is still open, with CI green, awaiting review.
Of the related contributions, PR #934 merged on August 13, #941 on the 17th and #942 on the 21st.
PR #953 is awaiting review from the maintainers.
The memory trade-off for that same optimization is still unmeasured: the new version keeps four columns of every page in the tree in memory while it builds the index. The net effect on memory usage has not been quantified yet.
Issue #14425 in Wagtail core, about the same bug in copy_for_translation(), is still open.
The demo does not yet show much of what wagtail-localize does: it lacks a language switcher, translated content in the initial data, and translatable snippets. That is left as follow-up work.
Issue #932 (refresh_segments()) is still open. It is the next step to bring in, following the same process as #953: use the harness to establish a baseline, the attribution script to confirm where the cost grows, and a reviewed PR before proposing it as a fix.
I want to thank the whole Wagtail team for an experience that has meant a great deal for my growth as a programmer. One of the most rewarding parts has been sharing these months with my mentors and GSoC peers. Especially Thibaud, from whom I've learned a lot, Meagen for her support throughout the whole program, and Coen for his guidance during the project.
I've learned from every meeting and conversation, not just from a technical standpoint, but also from the generosity and closeness with which this community works.
I also want to thank Google Summer of Code for making this opportunity possible and for its support of open source communities.