Skip to content

Instantly share code, notes, and snippets.

@rypptc
Last active August 23, 2026 18:34
Show Gist options
  • Select an option

  • Save rypptc/f5dbcd990bbadbd389f244a6e4db459a to your computer and use it in GitHub Desktop.

Select an option

Save rypptc/f5dbcd990bbadbd389f244a6e4db459a to your computer and use it in GitHub Desktop.
Scaling translation operations in wagtail-localize — GSoC 2026 final report

Student: Rodrigo Yáñez Pilgrim
Organization: Wagtail (Google Summer of Code 2026)
Mentors: Thibaud Colas, Coen van der Kamp
Project: Scaling translation operations in wagtail-localize
Project page: https://summerofcode.withgoogle.com/programs/2026/projects/wmfuUxir

Project description

The main goal of the project is to identify and optimize processes within wagtail-localize that have scalability problems, that is, processes whose cost grows with the size of the site. As a starting point, a preliminary investigation was carried out that identified several areas with potential for optimization.

The areas identified initially were the following:

  • Building the page index
  • The scalability of translation workflows, including processing page subtrees, saving translations, and moving the heaviest operations outside the HTTP request cycle
  • Detecting content changes before updating a translation

Redefining the project's scope

During the first two weeks of the project, the proposed flows were diagnosed. The result led us to rethink the project's scope, since in some cases the optimization was possible, but the benefit did not make up for the cost in code readability and maintainability.

A new investigation was carried out, which consisted of building wl-benchmark, a Wagtail project populated with enough volume of pages, articles, snippets, StreamFields and related objects, in multiple languages, to reproduce how a real site behaves at scale. This project was initially created locally as a testing lab. On this site, the extension's main flows were run, such as translating pages, subtrees and snippets, editing sources, syncing, and re-syncing the page tree across locales, measuring how many queries each operation made and how that number grew as the content size increased.

The measurement confirmed the page index problem that was already suspected: the cost grew with the size of the page tree, and a prototype showed it could be reduced to a fixed number of queries, regardless of how many pages there were. That area stayed unchanged in scope.

The other two areas did not have the same result. Content change detection was discarded. The method that was going to be optimized is barely used in the package's real flows — it mostly shows up in tests. The problem did not exist in practice.

The scalability of translation workflows had a mixed result. Grouping the save with bulk_create() was viable, but saving the translation for one locale already made few queries. It complicated the code without a clear benefit, so it was discarded. Moving save_target() outside the request cycle remains unresolved. It is not a query-count problem, it is a response-time problem, and that requires a different kind of measurement.

Looking closely at this area — the processing of subtrees and the full create_translations() flow — turned up a finding that was not on the original list: refresh_segments(), where every text segment extracted from a page fires several chained queries, with a cost that grows fast on pages with many segments. Alongside it appeared a smaller pattern, repeated in two different places: when processing related objects inside create_translations(), and when loading related segments inside save_target().

wl-benchmark results summary

The pattern confirms what had already been laid out in the scope redefinition: the page index and the related-object fan-out are resolved almost completely (they end up at a fixed number of queries, regardless of size); refresh_segments improves notably but does not reach a fixed number; and save_target and get_or_copy_target barely improve, confirming that they were not the right targets.

Raw output from wl-benchmark: each flow against its own database, freshly seeded with standard content (seed_realistic_localize_universe, default values), with nothing left over from a previous run.

Page index — synctree PageIndex scaling
synctree PageIndex scaling
                                 10 articles   50 articles  100 articles  250 articles
--------------------------------------------------------------------------------------
indexed pages                             53           105           170           363
aliases                                    2             8            15            36

baseline
  total queries                          260           520           845          1810
  from_database queries                  260           520           845          1810
  entry calls                             53           105           170           363
  duration                            63.6ms       129.6ms       212.4ms       543.3ms

prototype
  total queries                            2             2             2             2
  from_database queries                    2             2             2             2
  entry calls                             53           105           170           363
  duration                             6.3ms        10.9ms        18.0ms        39.6ms

reduction (total queries)
                                       99.2%         99.6%         99.8%         99.9%

Wrote profiler events: .tmp/benchmarks/synctree-index-scaling.jsonl
refresh_segments — refresh_segments scaling
refresh_segments scaling
                                  1 groups    2 groups    5 groups   10 groups   20 groups
------------------------------------------------------------------------------------------
segments                                27          43          91         171         331

baseline cold
  cold queries                         327         464         944        1716        3281
  cold duration                     59.4ms      82.5ms     168.8ms     294.0ms     550.2ms

prototype cold
  cold queries                         100         130         220         370         670
  cold duration                     34.1ms      46.6ms      87.4ms     149.0ms     290.6ms

baseline refresh
  refresh queries                      218         364         802        1532        2992
  segment queries                      187         333         771        1501        2961
  max segment queries                   10          10          10          10          10
  refresh duration                  44.8ms      72.7ms     149.5ms     275.7ms     549.1ms

prototype refresh
  refresh queries                       73         103         193         343         643
  segment queries                       32          62         152         302         602
  max segment queries                    2           2           2           2           2
  refresh duration                  30.5ms      48.1ms     142.3ms     150.0ms     268.9ms

reduction (refresh queries)
                                     66.5%       71.7%       75.9%       77.6%       78.5%

Wrote profiler events: .tmp/benchmarks/refresh-segments-scaling.jsonl
save_target — save_target scaling
save_target scaling
                                  1 groups    5 groups   10 groups   20 groups   50 groups
------------------------------------------------------------------------------------------

baseline (cold)
  cold queries                         211         221         231         251         311
  cold duration                     78.0ms      89.1ms     100.5ms     128.5ms     220.7ms

prototype (cold)
  cold queries                         209         209         209         209         209
  cold duration                     78.9ms      91.0ms      96.9ms     141.5ms     188.7ms

baseline (warm)
  warm queries                         159         167         177         197         257
  warm duration                     52.5ms      58.3ms      77.7ms     103.8ms     175.7ms

prototype (warm)
  warm queries                         155         155         155         155         155
  warm duration                     57.7ms     126.1ms      70.1ms     101.1ms     152.7ms

baseline (updated)
  updated queries                      162         170         180         200         260
  updated duration                 123.4ms      68.7ms      73.8ms     122.5ms     214.0ms

prototype (updated)
  updated queries                      158         158         158         158         158
  updated duration                  55.0ms      71.6ms      69.5ms     100.9ms     150.4ms

reduction (cold queries)
                                      0.9%        5.4%        9.5%       16.7%       32.8%

Wrote profiler events: .tmp/benchmarks/save-target-scaling.jsonl
Related-object fan-out — related object fan-out scaling
related object fan-out scaling
                                  1 groups    5 groups   10 groups   20 groups   50 groups
------------------------------------------------------------------------------------------
related segments                         3           7          12          22          52

baseline
  total queries                        128          30          45          75         165
  related queries                      126          28          43          73         163
  related calls                          3           3           3           3           3
  create calls                           3           3           3           3           3
  duration                          28.1ms       7.1ms      10.0ms      14.8ms      32.2ms

prototype
  total queries                         11          11          11          11          11
  related queries                        9           9           9           9           9
  related calls                          3           3           3           3           3
  create calls                           3           3           3           3           3
  duration                           3.7ms       3.6ms       3.8ms       4.4ms       4.4ms

reduction (total queries)
                                     91.4%       63.3%       75.6%       85.3%       93.3%

Wrote profiler events: .tmp/benchmarks/related-object-fanout-scaling.jsonl
get_or_copy_target — get_or_copy_target scaling
get_or_copy_target scaling

parents-ready
                                  1 groups   10 groups   50 groups
------------------------------------------------------------------

baseline
  total queries                        211         229         309
  copy queries                          66          66          66
  duration                          68.5ms      93.3ms     219.0ms

prototype
  total queries                        207         207         207
  copy queries                          66          66          66
  duration                          71.0ms      88.7ms     232.5ms

  reduction (total queries)
                                      1.9%        9.6%       33.0%

parents-missing
                                  1 groups   10 groups   50 groups
------------------------------------------------------------------

baseline
  total queries                        294         312         392
  copy queries                         149         149         149
  duration                          98.8ms     117.4ms     245.6ms

prototype
  total queries                        290         290         290
  copy queries                         149         149         149
  duration                          88.4ms     112.7ms     218.0ms

  reduction (total queries)
                                      1.4%        7.1%       26.0%

Wrote profiler events: .tmp/benchmarks/get-or-copy-target-scaling.jsonl

What was done

The benchmark harness

After seeing the results from wl-benchmark, it was proposed to bring this tool into the project on a permanent basis. Several measurement tools were suggested (query-doctor, Silk, django-eagle, django-o11y, nplusone, the pytest family). After analyzing them individually, django-query-doctor was identified as the most useful one for our case.

The harness was designed as a combination of custom tooling together with django-query-doctor. It includes a selection of wagtail-localize flows that showed signs of potential optimizations. In a first phase, it included 8 flows covering 15 executions. While trying to implement the page index case, the need to obtain query attribution came up, so an additional script was created for that.

The harness runs the package's real translation operations — submitting a page for translation, translating a snippet, syncing the page tree, refreshing segments, among others — in an isolated process, against a clean database created for each repetition, without reusing state between runs. It measures query count and time at different content sizes. Each run writes a JSON report with the commit, branch, and versions used, so any number can be reproduced. Queries are the primary result; time is secondary, since it depends on the machine it runs on.

Page index optimization

This optimization was initially proposed as a standalone PR (#930), designed based on the measurements taken with wl-benchmark. After the harness was implemented, it was decided to open a new PR built around the integrated tool, so the optimization would have a clear history that would make it possible to understand and reproduce the measurements in the future.

PR #953 runs the harness's three main scripts to get the diagnosis. PageIndex.from_database() cost 5 queries for every indexed page: the parent, the content type, the locale, and one query for each list of locales, the existing ones and the aliased ones. The cost grew linearly with the size of the tree — 119 queries for 24 indexed pages, 319 for 64.

The attribution script pointed to the 5 queries responsible for the growth, each at its line in synctree.py. The fix builds three mappings in a single pass and preloads the two missing relations. Result: from 119 to 2 queries at the small size, from 319 to 2 at the large size. Constant, regardless of the size of the tree.

When compared against query-doctor, the tool only detected 2 of the 5 real sources. The other 3 vary the translation_key on every call, so the query text is never identical and its duplicate detection does not group them. This is the same limitation that had already been anticipated when evaluating the tools.

The PR adds a regression test that checks the query count does not grow as the tree expands. The full suite, 608 tests, passes. One trade-off remains unmeasured: the new version keeps four columns of every page in the tree in memory while it builds the index — the net effect on memory was not measured, and remains as follow-up work.

Results

Before After
Queries at 24 indexed pages 119 2
Queries at 64 indexed pages 319 2
Growth per page 5.00 None
Median time at 64 pages 63.3 ms 2.8 ms
Raw output from the three commands (before the fix)
$ python benchmarks/run.py core_page_index --repeat 5
core_page_index [small]  119 queries  24.2 ms median (min 23.6, max 25.3, 5 repeats)  24 indexed_pages (expected 24)
core_page_index [large]  319 queries  63.3 ms median (min 62.0, max 66.4, 5 repeats)  64 indexed_pages (expected 64)
$ python benchmarks/run_attribution.py core_page_index
core_page_index: 24 -> 64 indexed_pages
  small   119 queries  = 119 attributed + 0 outside execute_wrapper
  large   319 queries  = 319 attributed + 0 outside execute_wrapper

  Growth drivers (5 of 5 growing groups)
     +1.00/indexed_pages     22 -> 62     synctree.py:72 in from_page_instance  (fixed -2)
             SELECT "wagtailcore_page"."id", "wagtailcore_page"."path", ...
     +1.00/indexed_pages     24 -> 64     synctree.py:75 in from_page_instance
             SELECT "django_content_type"."id", "django_content_type"."app_label", ...
     +1.00/indexed_pages     24 -> 64     synctree.py:77 in from_page_instance
             SELECT "wagtailcore_locale"."id", "wagtailcore_locale"."language_code" ...
     +1.00/indexed_pages     24 -> 64     synctree.py:79 in from_page_instance
             SELECT "wagtailcore_page"."locale_id" AS "locale" ... WHERE alias_of_id IS NULL
     +1.00/indexed_pages     24 -> 64     synctree.py:85 in from_page_instance
             SELECT "wagtailcore_page"."locale_id" AS "locale" ... WHERE alias_of_id IS NOT NULL

  Largest fixed cost (1 of 1 constant groups)
         1 fixed       1 -> 1      synctree.py:152 in from_database
             SELECT "wagtailcore_page"."id", "wagtailcore_page"."path", ...
$ python benchmarks/run_query_doctor.py core_page_index --size small
core_page_index [small]  24 indexed_pages
  119 queries, 2.5 ms of database time under diagnosis
  25 prescription(s)

  CRITICAL  n_plus_one  (24 queries)
    N+1 detected: 24 queries for table "django_content_type" (field: content_type)
    fix: Add .select_related('content_type') to your queryset
    at src/wagtail_localize/synctree.py:75 in from_page_instance

  CRITICAL  n_plus_one  (24 queries)
    N+1 detected: 24 queries for table "wagtailcore_locale" (field: locale)
    fix: Add .select_related('locale') to your queryset
    at src/wagtail_localize/synctree.py:77 in from_page_instance

The remaining 23 prescriptions are 22 duplicate_query warnings of two to four queries each, and one informational note about a missing index on depth in Wagtail's Page model.

Raw output from the three commands (after the fix)
$ python benchmarks/run.py core_page_index --repeat 5
core_page_index [small]  2 queries  1.6 ms median (min 1.5, max 1.7, 5 repeats)  24 indexed_pages (expected 24)
core_page_index [large]  2 queries  2.8 ms median (min 2.6, max 2.8, 5 repeats)  64 indexed_pages (expected 64)
$ python benchmarks/run_attribution.py core_page_index
core_page_index: 24 -> 64 indexed_pages
  small     2 queries  = 2 attributed + 0 outside execute_wrapper
  large     2 queries  = 2 attributed + 0 outside execute_wrapper

  Growth drivers (0 of 0 growing groups)
    none

  Largest fixed cost (2 of 2 constant groups)
         1 fixed       1 -> 1      synctree.py:158 in from_database
         1 fixed       1 -> 1      synctree.py:168 in from_database
$ python benchmarks/run_query_doctor.py core_page_index --size small
core_page_index [small]  24 indexed_pages
  2 queries, 0.2 ms of database time under diagnosis
  1 prescription(s)

The one remaining prescription is the missing index on depth, which is present before the change as well and belongs to Wagtail's own Page model.

Related contributions

Other contributions came up during the project. They do relate to the original proposal, but they belong to a broader view of the package.

The first was a demo site, as a reference implementation of wagtail-localize. It was proposed in issue #928 and opened as PR #929, the first of the program. Its review also raised whether the project should move to a src layout, so the PR was split in two: #941 for the layout and #942 for the demo.

When that first PR was published, a test failed during CI. It was investigated and found to be related to issue #922, open since before the program and unrelated to the demo: a bug in the translation of snippets with three or more levels of nesting. The fix was merged in PR #934. The same cause affects copy_for_translation() in Wagtail core, so it was reported there in issue #14425, which is still open.

The layout change was merged before the rest of the work. The harness and the page index optimization were already developed on top of the src layout and depend on the new paths, so that change had to land first.

Current status — what merged

The benchmark harness, PR #946, was merged on August 20, 2026.

The page index optimization went through two PRs. The first, #930, was designed with the wl-benchmark measurements, before the integrated harness existed. It was closed without merging on August 20, the same day the harness merged. It was replaced by #953, which redoes the same optimization on top of the now-finished harness. #953 is still open, with CI green, awaiting review.

Of the related contributions, PR #934 merged on August 13, #941 on the 17th and #942 on the 21st.

What's left

PR #953 is awaiting review from the maintainers.

The memory trade-off for that same optimization is still unmeasured: the new version keeps four columns of every page in the tree in memory while it builds the index. The net effect on memory usage has not been quantified yet.

Issue #14425 in Wagtail core, about the same bug in copy_for_translation(), is still open.

The demo does not yet show much of what wagtail-localize does: it lacks a language switcher, translated content in the initial data, and translatable snippets. That is left as follow-up work.

Issue #932 (refresh_segments()) is still open. It is the next step to bring in, following the same process as #953: use the harness to establish a baseline, the attribution script to confirm where the cost grows, and a reviewed PR before proposing it as a fix.

Acknowledgements

I want to thank the whole Wagtail team for an experience that has meant a great deal for my growth as a programmer. One of the most rewarding parts has been sharing these months with my mentors and GSoC peers. Especially Thibaud, from whom I've learned a lot, Meagen for her support throughout the whole program, and Coen for his guidance during the project.

I've learned from every meeting and conversation, not just from a technical standpoint, but also from the generosity and closeness with which this community works.

I also want to thank Google Summer of Code for making this opportunity possible and for its support of open source communities.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment