Skip to main content

A Cache Hit Is Not Free

Think a cache hit means the work is done? “A Cache Hit Is Not Free” offers a useful reminder: look past the hit and ask what it still costs before calling it a win.

Code Cracking
15m read
#Caching#Performance#Engineering
A Cache Hit Is Not Free - Featured blog post image
Mahmoud Zalt

1:1 Mentor

Are you a software engineer moving into AI?

Let's have a call. I'll help you modernize your skills and learn the tools, systems, and architecture behind reliable AI products. One session or ongoing.

Vibe Coding
with Confidence

The Vibecoder's Handbook, from idea to production

4.8

Everything you need to know about shipping software with AI, from the App idea to production.

What it covers

  • 0IntroductionWhat this book is & how to read it
  • 1Set UpGet your tools and a running app ready
  • 2PlanStructure your idea into a clear specification
  • 3ArchitectLay out a modular codebase for your AI
Start Reading Free

Triton compiler architecture

A persistent compilation cache removes lowering work, not the work required to hydrate and launch a kernel.

We are examining how Triton manages cached GPU compilation in compiler.py. Triton compiles Python-like GPU kernels into target-specific artifacts; this file coordinates cache lookup, backend compilation, artifact persistence, and the handoff to the GPU runtime. Its central object, CompiledKernel, matters because it separates “we have compilation output” from “this kernel can run on this device.”

That separation exposes an easily missed fact: a cache hit skips expensive lowering, but it still reconstructs state, reads artifacts, and defers GPU work until first use. I'm Mahmoud Zalt, an AI solutions architect, and we will trace those boundaries to learn the primary lesson: a cache hit is a lifecycle milestone, not a claim that execution-ready work is complete. We will follow the path from cache identity to hydration and first launch, then turn that model into design and operational guidance.

Two Lifecycles, One Handoff

compiler.py coordinates two distinct pipelines. The first produces or retrieves compilation artifacts. The second converts those artifacts into a GPU module that can launch. ASTSource and IRSource normalize different inputs behind a shared interface: each can provide a hash, produce IR, and contribute compilation options.

ASTSource or IRSource
          |
          v
      compile()
       /     \
 cache hit   cache miss
     |           |
     |     backend stages
     |           |
     +----- cached metadata + artifacts
                         |
                         v
                  CompiledKernel
                         |
                   first access
                         |
                         v
              validate + load binary
                         |
                         v
                       launch
Compilation and device initialization are separate lifecycles joined by CompiledKernel.

compile selects one backend, derives a cache identity, checks stored metadata, and runs ordered backend stages only on a miss. Backends own lowering and binary details; cache managers own persistence; driver.active owns device operations. The coordinator's key responsibility is sequencing those owners correctly.

Cache Identity Is Correctness

Before discussing the remaining cost of a hit, we need to establish what a hit means. The cache key is a correctness boundary: if it omits an input that changes generated code, Triton can retrieve the wrong artifact.

ASTSource.hash orders signatures and constants before hashing them. Sorting prevents dictionary insertion order from changing the identity of equivalent inputs:

def hash(self):
    sorted_sig = [v for k, v in sorted(self.signature.items())]
    get_key = lambda x: x.cache_key if hasattr(x, 'cache_key') else str(x)
    constants_key = '-'.join([get_key(v) for k, v in sorted(self.constants.items())])
    key = f"{self.fn.cache_key}-{str(self.attrs)}-{sorted_sig}-{constants_key}"
    return hashlib.sha256(key.encode("utf-8")).hexdigest()
AST source identity, lines 69–74.

This source hash is only one input. compile also uses get_cache_key to incorporate the source, backend, parsed options, and cache-invalidating environment values. This is content-addressed caching: output is located through an identity derived from the inputs that produced it.

On a metadata hit, Triton returns a CompiledKernel instead of rebuilding IR or running backend stages:

if not always_compile and metadata_path is not None:
    # cache hit!
    res = CompiledKernel(src, metadata_group, hash)
    if compilation_listener:
        compilation_listener(
            src=src, metadata=res.metadata._asdict(),
            metadata_group=metadata_group,
            times=timer.end(), cache_hit=True,
        )
    return res
The cache-hit return path, lines 267–278.

The important wording is “returns a CompiledKernel.” The hit avoids compilation, but object construction begins the next phase of work.

The Deferred Bill: Hydration, Then Launch

The CompiledKernel constructor reads JSON metadata, rebuilds the cached GPUTarget, selects a backend, and reads every non-JSON artifact into self.asm. A hit therefore still performs filesystem reads and retains the loaded artifact bytes in memory. Its cost grows with artifact count and total artifact size.

Lifecycle pointWork avoidedWork still performed
Compilation cache hitIR construction, lowering, artifact writesMetadata read, backend reconstruction, cached artifact reads
CompiledKernel createdGPU module loadingArtifacts retained in self.asm
First launchRepeated handle creation after initializationDevice lookup, resource checks, launcher creation, binary load
Later launchesCompilation and module loadingStream resolution, optional metadata and hooks, launcher call

This eager artifact loading makes inspection convenient, while SASS disassembly is deferred until asm['sass'] is requested. The asymmetry matters: constructing many cached kernels can consume disk bandwidth and memory even when only the final cubin or hsaco is needed. A sharper boundary would preserve paths for non-binary artifacts and load their contents on demand, while making an explicit, workload-specific decision about eager loading of the launch binary.

First launch pays a second deferred cost. CompiledKernel._init_handles creates the launcher, checks resource limits, and loads the binary. A compiled binary can still be unrunnable on the active device because its shared-memory or thread requirements exceed available capacity.

device = driver.active.get_current_device()
self._run = driver.active.launcher_cls(self.src, self.metadata)
shared = getattr(self._run, "shared", self.metadata.shared)
max_shared = max_shared_mem(device)
if shared > max_shared:
    raise_(OutOfResources(shared, max_shared, "shared memory"))

self.module, self.function, self.n_regs, self.n_spills, self.n_max_threads = \
    driver.active.utils.load_binary(self.name, self.kernel, shared, device)
Condensed from the lazy initialization path, lines 463–489.

That makes first launch a capability checkpoint, not merely a function call. The full path also validates thread limits and, when metadata provides it, tensor-memory limits. The file's architecture-specific tensor-memory constants reinforce an ownership rule: device capabilities belong with the driver or backend, rather than central compiler orchestration.

Design and Operate the Boundaries

Once we name cache lookup, artifact hydration, and device initialization as separate phases, improvement work becomes concrete. We can reduce work at the right boundary instead of treating “compile time” as one opaque number.

First, test cache identity as rigorously as compilation itself. Every source, target, option, backend-state, or relevant environment input that can change the binary must affect the key or invalidate the cache. At scale, concurrent same-key misses also deserve attention: this file has no lock coordinating them, so protection against duplicate compilation depends on cache-manager atomicity.

Second, validate inputs where their meaning is known. IRSource assumes its PTX regular expression found an entry prototype and calls match.group(1). A guard would turn an indirect AttributeError into a useful source error:

 if self.ext == "ptx":
     match = re.search(prototype_pattern[self.ext], self.src, re.MULTILINE)
+    if match is None:
+        raise ValueError(f"Unable to find a PTX entry prototype in {self.path}")
     self.name = match.group(1)
Validation belongs at the boundary that understands PTX input.

Finally, measure the transitions rather than only the public call. Cache-hit ratio tells us whether repeated workloads reuse results. Stage-level compilation time identifies lowering regressions. Creation-to-first-successful-launch time exposes hydration and driver-load latency that compile timing hides. These measures map directly to real lifecycle boundaries, so an alert can indicate which phase needs investigation.

  • Lookup: track cache-hit ratio by backend and target.
  • Lowering: record duration by backend stage.
  • Readiness: track creation-to-first-successful-launch time per GPU architecture.

That same boundary discipline supports production safeguards already present in the lifecycle: _module_pid avoids unloading a parent-owned GPU module in a forked child, and retaining a deep-copied launch failure avoids keeping traceback locals alive through a cached exception.

The Takeaway

The primary lesson is simple: a cache hit is not free because it establishes a new lifecycle boundary rather than completing the journey to execution. Triton correctly decouples artifact reuse from GPU initialization, but that design leaves observable work in hydration, validation, binary loading, and first launch.

We proved that by following the actual handoff: a content-addressed key returns CompiledKernel; its constructor reads and retains artifacts; its first device use validates resources and loads the binary. Each phase has a different owner, cost profile, and failure mode.

  • Audit what a hit still does. Follow the returned object's constructor and first-use path, not just the cache branch.
  • Model and measure explicit states. Separate cached, hydrated, device-loaded, and launchable in tests, telemetry, and latency budgets.
  • Keep policy with its owner. Put device capabilities in driver or backend layers, and keep orchestration focused on ordering lifecycle transitions.

As compiler infrastructure scales, the next optimization is rarely “make the cache branch shorter.” It is usually making the deferred boundary visible enough to choose what should be eager, lazy, deduplicated, and measured.

Full Source Code

Direct source from the upstream repository. Preview it inline or open it on GitHub.

heads/main/python/triton/compiler/compiler.py

triton-lang/triton • refs

Read Code on GitHub

Thanks for reading! I hope this was useful. If you have questions or thoughts, feel free to reach out.

Content Creation Process: This article was generated via a semi-automated workflow using AI tools. I prepared the strategic framework, including specific prompts and data sources. From there, the automation system conducted the research, analysis, and writing. The content passed through automated verification steps before being finalized and published without manual intervention.

Mahmoud Zalt

About the Author

I’m Zalt, a technologist with 16+ years of experience, passionate about designing and building AI systems that move us closer to a world where machines handle everything and humans reclaim wonder.

Let's connect if you're working on interesting AI projects, looking for technical advice or want to discuss anything.

Support this content

Share this article

Stay in touch

An occasional note when I build or write something new. Leave anytime.

Hire AI Employees

Hire AI Employees that work 24/7. No code.