Skip to content

Instantly share code, notes, and snippets.

@OhadRubin
Created April 20, 2026 09:28
Show Gist options
  • Select an option

  • Save OhadRubin/d080871ece839926cdc3fd444f5f353a to your computer and use it in GitHub Desktop.

Select an option

Save OhadRubin/d080871ece839926cdc3fd444f5f353a to your computer and use it in GitHub Desktop.
A survey of caching stances across reference systems

A survey of caching stances across reference systems

Nine entries, each from a different family. Close cousins are collapsed: React Query stands in for SWR and Relay, Nix stands in for Bazel and Buck, Dagster stands in for Luigi and Prefect and Airflow. Content-addressed storage (Git, OCI, IPFS) is folded into Nix because for a subsystem designer it's the derivation stance that's instructive; pure object dedup is a weaker commitment.


1. Spark — lineage-based recovery with opt-in materialization (also Flink, Beam)

What you reach for: df.cache() or df.persist(StorageLevel.MEMORY_AND_DISK).

Unit cached: a materialized DataFrame's partitions in memory/disk.

Key derivation: the logical plan. Spark's Catalyst compares query plans for equivalence — two calls that compile to the same plan can share a cached result.

Opt-in or opt-out: opt-in for materialization, opt-out for lineage (always on).

Stance: the lineage is the truth — the DAG of transformations from immutable source is always sufficient to reconstruct any intermediate. cache() is a performance optimization on top of that, trading memory for recomputation. Losing a cache entry (executor dies, partition evicted) is not an error; it's a trigger for recompute. This is why Spark caching has a different emotional temperature than, say, Redis caching — a miss is never user-visible as a failure. Defaults: be opt-in because blanket caching wastes memory; key on logical-plan equivalence, not object identity; fall back to silent recompute when an entry is lost.


2. tf.data — cache as a pipeline operator

What you reach for: dataset = dataset.map(parse).cache().shuffle(1000).batch(32); optionally dataset.cache("/tmp/scan.cache") to go to disk.

Unit cached: the materialized output of the pipeline up to this operator — a finite sequence of elements.

Key derivation: positional, within the dataset instance. In-memory caches are scoped to this dataset object in this process; on-disk caches key on the filepath string the user provides. There's no content hashing of the pipeline.

Opt-in or opt-out: opt-in per pipeline, and the position of the .cache() call is a design decision (after expensive parsing, before expensive shuffling).

Stance: caching is a node in a streaming graph. It exists to let epoch 2 skip the work epoch 1 already did, not to survive process restarts. The caller owns the position and the persistence decision. Defaults: place cache between the stable and variable parts of a pipeline; never try to detect "same dataset" across runs unless a path is provided; size-warn or fail on exhaustion, because the cache needs to materialize the whole scan before it can replay.


3. PyTorch torch.save(state_dict) — eager user-driven snapshot (also Flax checkpoints, Keras model.save)

What you reach for: torch.save(model.state_dict(), "ckpt_step_1000.pt") and later model.load_state_dict(torch.load("ckpt_step_1000.pt")).

Unit cached: a pickled dict of tensors — weights, optimizer state, epoch counter, anything the user chose to include.

Key derivation: the filepath the user typed. That's it. No content hash, no step-number convention enforced by the framework.

Opt-in or opt-out: fully opt-in. Nothing is ever saved unless torch.save is called.

Stance: the framework has no opinion about when to snapshot, how often, or where. The "cache" is a user-named blob written at user-chosen moments. It's not really a cache — it's persistence for transfer and resumability, and the framework is deliberately thin so users can compose their own policy (every N steps, on best-val-loss, before distributed sync, etc.). Defaults: never auto-save; let the user name the blob; round-trip save and load without loss; do nothing about invalidation because there's nothing to invalidate — old checkpoints are just old files.


4. Dagster IO managers — persistent output per asset (also Luigi targets, Prefect results)

What you reach for: writing @asset(io_manager_key="s3_pickle") on the computation; configuring the io manager once per environment. In Luigi: returning luigi.LocalTarget("path/to/output.csv") from Task.output().

Unit cached: the output of a task or asset — a dataframe, a model, a file, whatever the computation returned.

Key derivation: the asset key plus partition key, deterministically mapped to a backing store location by the IO manager. Dagster tracks materialization timestamps; Luigi uses target existence as the complete signal.

Opt-in or opt-out: opt-out in practice. Every asset has a backing store by default; skipping persistence is the exceptional case.

Stance: the cache is indistinguishable from the output. Every node in the DAG is durable by default; re-running a task means "check if its target exists; if so, skip." Invalidation is manual (delete the target, or invoke --re-execute), time-based (partition expiry), or metadata-driven (upstream asset newer than mine). The asset's presence in the backing store is the entire cache-hit signal. Defaults: every node persists; skip re-runs when the target is there; invalidate by deletion or explicit re-execution rather than by hashing inputs; make every cached artifact a thing you can point at with a URI.


5. joblib.Memory — content-hashed memoization to disk

What you reach for: memory = Memory("./cachedir"); @memory.cache\ndef train(X, y, lr): ....

Unit cached: the return value of a decorated function, pickled to disk.

Key derivation: a hash over the function's source code plus a content hash of the arguments (including numpy arrays, pandas DataFrames, custom classes via their pickle bytes). This is the pivotal move — joblib hashes tensor content, not object identity, so a freshly-loaded numpy array is a cache hit against a previously-seen identical array.

Opt-in or opt-out: opt-in per function.

Stance: memoization where "same inputs" means "same bytes," not "same reference." Code changes to the function body invalidate the entry automatically because the source is part of the hash. Durable across process restarts because it's on disk. This is the stance that makes research workflows survive Jupyter kernel restarts without recomputing the data-loading step. Defaults: hash by content of arguments, not identity; include the function source in the key so edits invalidate; store on disk with GC by age/size; fail loudly when an argument isn't hashable rather than silently skipping the cache.


6. Nix / Bazel — content-addressed derivation

What you reach for: writing a derivation in default.nix (stdenv.mkDerivation { src = ./src; buildInputs = [ gcc openssl ]; }) or a BUILD rule. The cache surface is implicit in the declaration; you never call cache().

Unit cached: build outputs — files, directories, compiled artifacts, whole filesystem trees.

Key derivation: a content hash of the full transitive closure of inputs. For Nix: sources, build script, every dependency recursively, the compiler, the standard environment. Bazel: the action hash (inputs + command + env). Same hash means same output, guaranteed by construction.

Opt-in or opt-out: opt-out and you can't really opt out. Everything is cached; there's no "just build without checking the cache" path.

Stance: caching is inseparable from computation. The cache isn't a speed optimization — it's the system's model of correctness. An old derivation hash resolves to the same bytes forever, across machines, across years; "stale" has no meaning in this world. A remote cache isn't a risk, it's a truth-propagation mechanism. Defaults: never invalidate an entry; hash everything that could possibly affect output including the toolchain; treat a cache hit as a proof, not a guess; make it easy to share caches across machines because the hash guarantees safety.


7. dbt incremental — warehouse-state materialization (also Postgres materialized views)

What you reach for: {{ config(materialized='incremental', unique_key='id') }} at the top of my_model.sql, plus an {% if is_incremental() %} block that filters the source to only rows newer than what's already in the table. In Postgres: CREATE MATERIALIZED VIEW foo AS SELECT ...; REFRESH MATERIALIZED VIEW foo;.

Unit cached: a SQL table in the warehouse — actual persistent rows.

Key derivation: the model name maps to a table name. There's no content-based identity at all; freshness is either a timestamp column, a high-watermark, or wall-clock refresh cadence.

Opt-in or opt-out: the materialization strategy is user-chosen per model — view (no cache, recompute on every read), table (full rebuild each dbt run), incremental (append/merge new rows). The choice is the entire design.

Stance: the warehouse IS the cache. There's no separate cache store; the production table is both output and cached intermediate. Incremental logic is essentially a manual edit-distance model: given "the table as of last run" and "new source rows," produce "the updated table." This stance makes caching a first-class part of the data model, not a system-level concern. Defaults: pick materialization per model based on recomputation cost vs freshness tolerance; make --full-refresh a cheap ritual; invalidate by wall-clock, high-watermark, or explicit refresh.


8. HTTP cache-control + eTags — origin-declared freshness contract

What you reach for: on the origin server, headers — Cache-Control: public, max-age=3600, stale-while-revalidate=86400, ETag: "v3-abc". On the client, nothing explicit; the cache layer reads the headers and acts.

Unit cached: HTTP response bodies, with their headers as metadata.

Key derivation: (URL, method, varying request headers as declared by Vary). Revalidation uses If-None-Match: "v3-abc" against the eTag, or If-Modified-Since.

Opt-in or opt-out: the origin opts caches in or out via headers. Clients can override with Cache-Control: no-cache in their requests but mostly don't.

Stance: caching is negotiated between origin and intermediaries via a rich vocabulary. The origin is the authority on freshness; the cache is a faithful executor of the contract. stale-while-revalidate, must-revalidate, private vs public, immutable, no-store — all are first-class words in the protocol. Defaults: let the origin declare; respect the directives; revalidate cheaply with eTags rather than refetching; split the cache on every Vary dimension; assume missing headers mean "don't cache."


9. React Query / SWR — stale-while-revalidate as a UI primitive

What you reach for: const { data, isLoading, isStale } = useQuery(['users', id], () => fetchUser(id)). Invalidation: queryClient.invalidateQueries({ queryKey: ['users'] }).

Unit cached: the result of a fetcher function, indexed by a user-provided query key tuple.

Key derivation: the query key tuple, with deep structural equality. ['users', 7] and ['users', 7] are the same entry; ['users', 7, 'posts'] is a different entry that invalidation can target by prefix.

Opt-in or opt-out: opt-out within the hook — every useQuery caches automatically; disabling is the special case.

Stance: the cache is the component tree's local model of remote state. "Stale" is a first-class state that coexists with "fetching" rather than replacing it — you render stale data immediately and let the background refetch update the UI when it finishes. Invalidation is semantic (invalidate this query-key prefix) rather than TTL-centric, though TTLs exist. This stance is incompatible with treating the cache as invisible; the cache is the component's data source. Defaults: show stale data immediately, refetch on focus and reconnect, dedupe in-flight requests by key, expose staleness separately from loading, scope invalidation by key prefix.


The axes along which these genuinely disagree

Axis 1 — Why does caching exist: to avoid recomputation, or to guarantee determinism?

  • Recomputation-cost driven: Spark .cache(), tf.data, joblib, React Query, HTTP (mostly)
  • Determinism-driven: Nix/Bazel (cache hit = proof of equivalence), joblib's source-hash component
  • Transfer / resumability driven: torch.save, Dagster assets, dbt tables

This is the axis that catches the most expensive mistakes. A designer who thinks their cache is about speed but has really built a determinism contract (by content-hashing inputs, say) will feel betrayed when "refresh the cache" makes no sense as an operation. Going the other way, a designer who thinks they're building a Nix-like truth store but is secretly TTL-based has built a bug farm.

Axis 2 — Do you store the value, or the recipe for producing it?

  • Value stored: joblib, Nix, dbt table, Dagster, torch.save, React Query, HTTP, tf.data on-disk
  • Recipe stored (cache is recomputation-from-lineage): Spark lineage, dbt view
  • Both, with value as an optimization over recipe: Spark .cache() (the lineage is always there; the materialization is an optimization)

The view vs table choice in dbt is this axis at a single decision point. Same file, opposite stances — and picking the wrong one is one of the more common dbt mistakes. This axis also explains why Spark's cache feels different from Dagster's assets even though both persist intermediate results: Spark treats persistence as secondary to lineage; Dagster treats persistence as primary.

Axis 3 — Whose job is invalidation?

  • The cache never invalidates, by construction: Nix
  • The cache invalidates when the source changes (automatic, content-hash-driven): joblib
  • The origin declares invalidation via protocol: HTTP
  • The user invalidates manually (delete the target / re-run with flag): Dagster, Luigi, torch.save, dbt
  • The system invalidates on lifecycle events (focus, reconnect): React Query
  • Not applicable — invalidation is bounded by lifetime: tf.data within a training run, Spark within a session

If the subsystem you're designing doesn't have a clear answer to "whose job is it to say this entry is wrong now," it probably hasn't decided its stance yet. Different traditions put the responsibility in wildly different places, and borrowing interface from one while borrowing invalidation semantics from another tends to produce the bug class called "stale data forever."

Axis 4 — Is the cache a transparent layer or a visible node you can point at?

  • Transparent (the cache is behind the API, you just call the function): joblib, HTTP, Nix (the store exists but you rarely interact with it directly), Spark .cache()
  • Opaque (the cache is a first-class thing in the system with a name, a location, a URI): Dagster assets, Luigi targets, dbt tables, torch.save files, tf.data on-disk caches, React Query's devtools-inspectable cache

This axis is largely about whether your cache entries are addressable by users directly. When they are, users expect to be able to point at them, re-run only downstream of them, inspect them, partially invalidate them — a whole surface of operations the transparent-layer tradition doesn't need to support. The orchestration family lives firmly on the opaque side; the memoization family firmly on the transparent side. A subsystem trying to be both usually ends up delivering neither.


Sentences to complete

  • Closest to Spark lineage → default to opt-in materialization, always retain the graph, recompute silently on cache miss.
  • Closest to tf.data's .cache() operator → default to cache-as-a-pipeline-step, scoped to one pipeline instance, requiring an explicit path for cross-run persistence.
  • Closest to torch.save → default to nothing-automatic, user names the blob, framework stays out of the policy.
  • Closest to Dagster/Luigi → default to every-output-persists, presence-means-hit, delete-to-invalidate, addressable URIs.
  • Closest to joblib.Memory → default to content-hash over arguments, source-hash the function, disk-backed, fail loud on unhashable inputs.
  • Closest to Nix/Bazel → default to content-hash everything, never invalidate, treat a hit as a proof, remote cache is safe by construction.
  • Closest to dbt incremental → default to making the storage be the cache, pick materialization per asset, invalidate by watermark or --full-refresh.
  • Closest to HTTP cache-control → default to origin declares, clients obey, revalidate cheaply, split on declared varying dimensions.
  • Closest to React Query → default to stale-is-usable, refetch on lifecycle events, semantic invalidation by key prefix, expose staleness separately from loading.

The easiest trap is landing between two of these without noticing — an interface borrowed from memoization on top of semantics borrowed from orchestration, or a persistence model borrowed from HTTP on top of an invalidation model borrowed from Nix. The test is whether you can finish the sentence unambiguously. If your subsystem requires an "and" — "closest to Dagster and joblib" — it's either a legitimate two-layer composition (persistent assets whose contents happen to be memoized, say), or it's a design that hasn't picked yet.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment