This page starts from the simplest matmul you can write in PyPTO and follows it to the code that runs on an Ascend AI Core. There is no hand-tuned kernel anywhere on the page. Every tile, buffer, loop, and synchronization flag you will see was put there by the compiler stack: PyPTO’s pass pipeline, the simpler runtime, and the PTOAS assembler.
@pl.jit def matmul(a: pl.Tensor[[128, 512], pl.FP16], b: pl.Tensor[[512, 128], pl.FP16], c: pl.Out[pl.Tensor[[128, 128], pl.FP32]]): with pl.at(level=pl.Level.CORE_GROUP): c[:] = pl.matmul(a, b, out_dtype=pl.FP32) return c
Part I follows that line down, stage by stage, and ends at the instruction stream. Part II asks the
question an HPC reader asks next: this kernel runs on one core. Which component makes it use
all of them, keep weights in the hardware’s native layout, or handle matrices that don’t fit
on the chip? For each of those, the answer is a specific pass, a specific runtime feature, or
“you, in one more line of DSL”.
Provenance. Structure is observed: every pass diff, buffer address,
instruction, synchronization flag and dependency edge below is copied from one compile of this
program and from the runtime’s own dependency dump. Target a2a3sim, backend
Ascend 910B, PyPTO ee49fcea. The simulator runs the real kernel (the result matches
PyTorch to 7.6×10−5) but is not timing-accurate, so nothing on this page is a
cycle count, a bandwidth, or a speedup.
Before the compiler runs
The program names no memory space, no tile size and no loop. It still commits to four things, and each one decides something later on the page.
c is a caller-owned buffer the program writes. That marking survives all the way to
the runtime, which uses it to decide that c needs no copy to the device before the run
and a, b need no copy back after it.Why placement must be explicit on this target
For this page’s Ascend 910B target (a 220x-class separation-mode device), the relevant execution
model is MPMD-like: the AIC (Cube) and AIV (Vector) cores have independent scalar
control paths and can load separate instruction streams. Data exchanged between them follows the
documented global-memory path on this target, while explicit cross-core flags provide ordering.
PyPTO therefore cannot rely on a transparent GPU-style cache hierarchy for tensor tiles; placement
and movement have to be exposed or inferred. A manual six-statement matmul makes that contract
visible: two GM→Mat loads, two Mat→Left/Right moves, one cube multiply producing Acc, and
one store back to GM. The whole-tensor example on this page is a separate route: the
AutoTileMatmulL0 pass chooses legal L0 tiles with a cost model, then PyPTO emits PTO
MLIR and PTOAS lowers it to C++ calls into PTO-ISA before compilation.
The map
PyPTO’s passes turn the program into two functions. The device kernel goes through PTOAS to C++ calls into PTO-ISA for the AI Core. The orchestration function becomes C++ against simpler’s API and runs on the AI CPU. simpler brings the two back together at run time. Step through once; each later section expands one box.
What the passes are placing tiles into
Before watching the compiler place tiles, look at where it can put them. Each on-chip memory space is its own SRAM with its own capacity, its own owning core, and a short list of legal ways in and out. They are separate buffers, not levels of a cache: L0A is not “inside” L1 the way a CPU’s L1 sits inside its L2. Click or tab to any space or arrow.
The whole pipeline, one line per pass
PyPTO lowers the program through 54 passes over one IR. Tensor ops and tile ops live in the same tree; pass 12 is where one turns into the other. Diffing every pass’s output against its input shows that 16 of them change this program. The table lists all 54: what each pass is for in general, and what it did here or why it had nothing to do. Highlighted rows changed the IR; click one to open its diff in the stepper below.
| # | Pass | What it does | This program | Diff |
|---|
Phases are this page’s grouping of the pass manager’s order. Pass numbers are the
pass manager’s invocation order, which is what passes_dump/NN_after_<Pass>.py
files are named after; the pass documentation numbers each one lower (this page’s pass 19 is doc 18,
pass 37 is doc 36). The pass list, order and
descriptions come from the compiler’s own pass registry and documentation; the “this
program” column comes from diffing consecutive IR dumps of this compile.
Diffs are the real IR text of consecutive dumps, shown as unified diffs with three
lines of context. For readability the dumps were run through ruff format (88 columns) before diffing, tile
layout annotations (pl.TileView(…)) are dropped, and pl.const(N, pl.INT64) is printed as
N. The +/− counts are of the displayed lines.
Who wrote Left and Right?
A pass called InferTileMemorySpace exists, and you would expect
it to place the operands. Here it doesn’t. Pass 12 already put A and B in Mat and the
result in Acc, and pass 19 wrote target_memory=Left/Right
explicitly on its extracts. It has to: it runs two passes before inference, at a point where every tile
is 2D but no space has been inferred yet. By pass 21 there is nothing left to infer.
Double buffering is three passes, not one
AutoTileMatmulL0 (19) only asks for two stages.
LowerPipelineLoops (32) copies the loop body so consecutive iterations touch different
buffers. CanonicalizeIOOrder (33) moves the second copy’s loads above the first
copy’s matmul, so the DMA engine has work queued while the cube computes. Then
MemoryReuse (37) is told not to merge the two copies’ buffers, which is the only
reason two of each survive to allocation.
The one real decision in the pipeline
Of the sixteen passes that changed something, one made a choice with alternatives:
AutoTileMatmulL0 picked the L0 tile. The operands as written cannot sit in the
cube’s operand buffers at all. A is 128 × 512 FP16, which is 128 KiB, and L0A holds 64 KiB.
Something has to split K.
FP16 operands (2 bytes), FP32 accumulator (4 bytes), Ascend 910B capacities. This is residency arithmetic only. It shows which shapes are legal, not which legal shape is fastest.
The rule that picks k = 128
When both operands stream through the cube (the output-stationary case, where the accumulator stays put and A and B slices flow past it), the chooser plans for two copies of each operand tile from the start. It sizes k against half of L0A: 64 KiB / (2 bytes × 2 stages) = 16 384 elements, so with m = 128 the largest legal k is 128. k = 256 would fit one buffer exactly and leave no room for the second. The pick fills L0A and L0B to exactly 100 %.
A search with a cost model, not a table
The chooser enumerates every legal 16-aligned (m, n, k),
crossed with which operand stays resident and whether the accumulator is double-buffered, and scores
each candidate with a roofline model of cube cycles, operand-load cycles (weighted by the unequal
L0A and L0B port bandwidths) and accumulator drain cycles. It returns the minimum. The model is
checked against a brute-force re-enumeration of itself, which proves the search is correct. Nobody
has timed the pick against alternatives on hardware.
The chosen 128 × 128 tile is not a hardware parameter. The cube works on fractal boxes, and
inside every on-chip buffer a tile is stored as a grid of them: a 512-byte box (16 rows of 32 bytes)
for the FP16 operand tiles, a 1024-byte box for the FP32 accumulator (visible as
fractal=512 and fractal=1024 in the compiled tile types). The tiler’s
constraints are stated in those units: every tile dimension a multiple of 16, and K 16-aligned or the
pass declines. 128 is just the largest multiple of 16 that satisfies the double-buffered budget above.
The chooser splits K before it ever splits M or N. Splitting K costs nothing extra: the partial sums accumulate in one L0C tile and drain to memory once. Splitting M or N means one drain per output sub-tile. So it tiles the output only when the output itself doesn’t fit L0C, and here a 128 × 128 FP32 accumulator is 64 KiB of a 128 KiB buffer, so it doesn’t.
The territory, as the allocator filled it
Each panel is one memory space, drawn to that space’s real capacity. Time runs downward through the kernel’s statements, so a block’s height is how long its buffer is live. Every space starts its own address line at zero, which is what “separate buffers, not a hierarchy” means in practice. Hover a block or a statement.
Full, on purpose
L0A and L0B are at 100 % and L0C at 50 %. That is the tile choice from the previous section, showing up as addresses: ping at offset 0, pong at 32 KiB, nothing left over. Mat holds all of A and B (256 KiB of 512) for the whole kernel, because the program fits on the chip. Part II shows what happens when it doesn’t.
The runtime half of the program
Pass 10 split the program in two, and so far the page has followed the kernel. The other half, the orchestration function, is also compiled. It becomes a small C++ library that runs on the device’s AI CPU, not on the host, and its only job is to tell simpler’s scheduler what tasks exist and which tensors each one reads and writes. This section follows that half from the compiler to the moment the cube starts.
Nothing about the runtime is decided in a separate tool. Four of the sixteen passes that changed the IR exist to shape the orchestration side, and each one maps to one line of the generated file below.
| Pass | IR change here | What it becomes in the orchestration file |
|---|---|---|
10 OutlineIncoreScopes | the pl.at body becomes kernel matmul_matmul; the caller keeps one call | one task, params_t0, and one submit call |
25 ExpandMixedKernel | kernel retyped InCore → AIC (pure cube, no vector half) | rt_submit_aic_task, not the mixed-cluster form |
42 DeriveCallDirections | arg_directions = [input, input, output_existing] | add_input, add_input, add_output |
51 MaterializeRuntimeScopes | body wrapped in one pl.scope() | one SIMPLER_SCOPE() { … } |
This is the whole file, as emitted. Click a line.
SIMPLER_SCOPE is not a barrier
It reads like a fork-join block, and it isn’t one. Closing a scope doesn’t wait for the scope’s tasks to finish. It drops the scope’s reference on each task it submitted, which is what allows the runtime to recycle those tasks’ ring slots and intermediate buffers once their consumers are done. Ordering between tasks comes from somewhere else entirely: the tensors each task reads and writes, as the next part shows.
From the Python call to the cube and back
A run involves three separately compiled programs: simpler’s host library, the orchestration
library on the AI CPU, and the kernel on the AI Core. Step through one call of
matmul(a, b, c).
With dependency capture on, the runtime writes down every task it was given and every edge it derived. For this program: one task on the cube, three tensor arguments, no edges. The direction tags are the ones pass 42 derived.
There is no edge list in the orchestration code. When a task is submitted, the runtime looks up each of its input tensors in a producer table (the TensorMap) and adds an edge from whichever earlier task last wrote it, then records this task as the producer of its outputs. A task with no producers goes straight onto the ready queue for its core type. So the direction tags from pass 42 aren’t metadata; they are the dependency graph.
That has consequences once there is more than one task. Part II has a case where a documented multi-core pattern quietly runs on one core at a time, and the reason is visible in exactly this kind of dump.
From tile IR to C++, and where the synchronization comes from
After pass 54 the kernel is printed as an MLIR module in the PTO dialect and handed to PTOAS, the
assembler. In that module the memory spaces are still part of each tile’s type
(loc=mat, loc=left, loc=acc), and so are the layouts from the
memory-space table. What the module does not contain is a single synchronization
instruction. PyPTO has no intra-core sync pass. PTOAS inserts every set_flag and
wait_flag the kernel needs, and it is the only component that does.
in — the loop body of the .pto module (types shortened)
scf.for %ko = 0 to 512 step 256 {
%a0 = pto.alloc_tile addr = 0 : tile_buf<loc=left, f16, 128x128>
pto.textract ins(%a_mat, 0, %ko) outs(%a0)
%b0 = pto.alloc_tile addr = 0 : tile_buf<loc=right, f16, 128x128>
pto.textract ins(%b_mat, %ko, 0) outs(%b0)
%a1 = pto.alloc_tile addr = 32768 : tile_buf<loc=left, …>
pto.textract ins(%a_mat, 0, %ko+128) outs(%a1)
%b1 = pto.alloc_tile addr = 32768 : tile_buf<loc=right, …>
pto.textract ins(%b_mat, %ko+128, 0) outs(%b1)
scf.if %ko == 0 { pto.tmatmul(%a0, %b0) }
else { pto.tmatmul.acc(%acc, %a0, %b0) }
scf.if %ko == -128 { pto.tmatmul(%a1, %b1) }
else { pto.tmatmul.acc(%acc, %a1, %b1) }
}
// no set_flag / wait_flag anywhere in the moduleout — the same loop in the emitted C++ (declarations dropped)
wait_flag(PIPE_MTE2, PIPE_MTE1, EVENT_ID0); // A, B in Mat
for (ko = 0; ko < 512; ko += 256) {
wait_flag(PIPE_M, PIPE_MTE1, EVENT_ID0); // ping free?
pipe_barrier(PIPE_MTE1);
TEXTRACT(a0, a_mat, 0, ko);
TEXTRACT(b0, b_mat, ko, 0);
set_flag(PIPE_MTE1, PIPE_M, EVENT_ID0); // ping full
wait_flag(PIPE_M, PIPE_MTE1, EVENT_ID1); // pong free?
TEXTRACT(a1, a_mat, 0, ko + 128);
TEXTRACT(b1, b_mat, ko + 128, 0);
set_flag(PIPE_MTE1, PIPE_M, EVENT_ID1); // pong full
wait_flag(PIPE_MTE1, PIPE_M, EVENT_ID0);
if (ko == 0) TMATMUL(acc, a0, b0);
else TMATMUL_ACC(acc, acc, a0, b0);
set_flag(PIPE_M, PIPE_MTE1, EVENT_ID0); // ping free
wait_flag(PIPE_MTE1, PIPE_M, EVENT_ID1);
if (ko == -128) TMATMUL(acc, a1, b1);
else TMATMUL_ACC(acc, acc, a1, b1);
set_flag(PIPE_M, PIPE_MTE1, EVENT_ID1); // pong free
}Both listings are excerpts of the real files from this compile, with long template types
and variable declarations removed and SSA names shortened (%t__tile_l0_a →
%a0). The comments on the right are this page’s.
Read the event IDs as buffer names
EVENT_ID0 is the ping buffer pair and EVENT_ID1 the pong pair.
Each pair has two flags going in opposite directions. MTE1→M says “this buffer is full,
compute may read it” (read-after-write). M→MTE1 says “compute has finished with this
buffer, the DMA may overwrite it” (write-after-read). Before the loop, PTOAS pre-sets both
M→MTE1 flags, so the first two extracts don’t wait for a matmul that never happened; after
the loop it drains them.
How PTOAS gets there
The sync pass runs six stages in order. It translates the module into
a list of operations tagged with the pipe each runs on and the buffers it touches, finds every
cross-pipe hazard, moves sync points to cheaper positions, removes
syncs that others already cover, assigns hardware event IDs, and emits the calls. The loop bound and
both ko == … tests come through unchanged; the second one can never be true,
because pass 49 left it unfolded.
| Instruction | Pipe | Hardware | Issued |
|---|---|---|---|
TLOAD × 2 | MTE2 | DMA: global memory → L1 | once, before the loop |
TEXTRACT × 4 per trip | MTE1 | DMA: L1 → L0A / L0B | 8 in total |
TMATMUL / TMATMUL_ACC | M | the cube | 1 + 3 in total |
TSTORE | FIX | FixPipe: L0C → global memory | once, after the loop |
What stage=2 actually buys
Each pipe (the two DMA engines, the cube, the FixPipe) executes its own instructions in order, and the
pipes run in parallel with each other. The only thing that holds one pipe back for another is a flag.
This figure places every instruction of the kernel at the earliest step its flags allow, assuming
every instruction takes one step. Switch to stage=1 to see the same kernel with a single
buffer pair.
What the figure shows
With two buffer pairs, the extract for iteration i+1 goes into the other half while matmul i is still reading the first, so MTE1 and the cube stay busy on the same steps. The dashed violet edges are the new constraint double buffering introduces: iteration i+2 reuses iteration i’s half, so its extract waits for matmul i to finish. The amber chain is the one dependency no buffering can remove: all four matmuls accumulate into the same L0C tile, in order.
What it does not show
Positions are dependency steps, not time. On real hardware a 128 × 128 × 128 FP16 matmul and a 32 KiB L1→L0 copy take different numbers of cycles, and whether MTE1 or the cube is the bottleneck depends on those numbers. The simulator this page used runs the kernel correctly but does not model timing. The step counts are a statement about the kernel’s dependency structure, which is exactly what the flags encode.
Part II
Part I’s kernel is correct, tiled for L0, double-buffered and synchronized, and it runs on one of the 910B’s 24 cube cores with operands that fit on the chip. Everything an HPC matmul adds beyond that has an owner. Here is who owns what, checked against the compiler, the runtime and the assembler source, and then each missing piece added to the same program and compiled.
| Optimization | Who does it | How | In Part I? |
|---|---|---|---|
| L0 tiling: K-loop, and M/N split when the output overflows L0C | PyPTO pass 19 | automatic cost-model search | yes: k = 128 |
| L0A/L0B ping-pong inside that loop | PyPTO passes 19, 32, 33, 37 | automatic once the matmul is tiled | yes: 2 × 2 buffers |
| Intra-core synchronization and event IDs | PTOAS sync pass | automatic always enabled by PyPTO | yes: 2 events |
| Fractal layout for tiles in Mat / L0 | PyPTO tile conversion | automatic converted during the load | yes |
| Accumulator (L0C) double buffering | PyPTO passes 19, 32, 33, 37 | automatic or opt-in automatic under memory_planner=PTOAS and DSA_RP; a marked-experimental flag under the default PyPTO planner | no — k = 128 is split-K, and dbC = 2 requires the full-K emitter |
| Using more than one core: M/N split or split-K | you, in the DSL; simpler dispatches | one more line pl.spmd(n) | no — rungs 1, 2 |
| L1 tiling: operands bigger than the 512 KiB Mat | nobody | compile error you block K yourself | no — rung 3 |
| Double buffering of your own loops (GM→L1) | you; PyPTO lowers it | one more line pl.pipeline(…, stage=2) | no — rung 3 |
| Weights stored in the hardware layout | you; PyPTO pass 16 plans for it | a type annotation pl.NZ | no — rung 4 |
| Epilogue fusion (bias, ReLU, quantize on writeback) | you, at tile level | tile-level store options | no |
| Overlap between tasks | simpler | automatic for tasks you created, ordered by their tensors | one task: nothing to overlap |
What those four passes each do
The ping-pong row names four passes because no single pass owns this optimization. Reading them in order is the whole mechanism — and the fourth one is the one that people usually assume is the second.
AutoTileMatmulL0 requests it. The tiled K-loop is emitted as
ForKind::Pipeline with pipeline_stages=2. That 2 is the
operand depth — L0A and L0B — not the accumulator. It is also why
k = 128 and not 512: the chooser sizes the tile against half of each operand buffer,
which is the arithmetic in the previous section.LowerPipelineLoops replicates the body. Two copies per outer iteration, each
with fresh def-vars, and each clone's tile-producing call tagged with a
pipeline_membership (group, stage) pair.CanonicalizeIOOrder shapes the schedule — scalars first, then the loads
clustered ahead of compute, so the load engine runs ahead of the cube. It only reorders statements; it
does not create the buffer separation.MemoryReuse is where the separation is actually enforced. The two clones are
sequential in program order, so their tiles have disjoint lifetimes — exactly the
condition the reuse packer coalesces on. Left alone it would fuse the ping-pong back into a single
buffer. The pipeline_membership tag is what forbids that, and it does so only as deep as
the memory space can afford.The accumulator is deliberately left untagged: it is written
by the serialized cube, one tile's MAD retiring before the next begins, so a single L0C buffer is already
correct. That is why this page shows one Acc @ 0 rather than a pair.
The pattern
The compiler automates the single-core, fits-on-chip story
completely. Everything about many cores and data bigger than the chip is written by
the programmer today, as small DSL additions to the same program. That split is deliberate: the
assembler’s design notes rule out inferring multi-buffering on its own, the accumulator double
buffer is experimental under the default planner (it is automatic under
memory_planner=PTOAS and DSA_RP, and needs the full-K emitter either way),
and nothing in the compiler splits one device region across cores.
Production kernels built on PyPTO write all four rungs below by hand.
The switch this page did not have to touch
Everything above was compiled with the default
memory_planner=PYPTO. That setting decides who owns on-chip memory, and it is the one knob
that can also move the tile the chooser picks.
PYPTO (default) | PTOAS | |
|---|---|---|
| Addresses | baked addr on every tile; ptoas runs at --pto-level=level3 | addr-less tiles, placed by ptoas at --pto-level=level2 |
| Reuse and placement | PyPTO MemoryReuse + AllocateMemoryAddr | PTOAS PlanMemory, with both of those skipped |
| L0C double buffer | opt-in flag | automatic, if the chooser picks it |
It can change the tile. dbC = 2 requires the full-K
emitter — l0_tile_chooser.cpp sets require_full_k = !is_os || r.dbc —
and halves the L0C accumulator budget (l0c_bytes / (bytes_c × 2)), so making it
available re-ranks the candidate shapes. On this program it changed nothing: the same
128 × 128, k = 128, one accumulator, at identical offsets, under all
three configurations. That is the useful part — Part I is a default-planner picture, and it holds
because k = 128 split-K won on modelled cost, not because a full-K
dbC = 2 candidate was illegal.
Rung 1 · more cores, split the output
pl.spmd# M = 512, N = 128, K = 512
@pl.jit
def matmul_spmd(a: pl.Tensor[[512, 512], pl.FP16],
b: pl.Tensor[[512, 128], pl.FP16],
c: pl.Out[pl.Tensor[[512, 128], pl.FP32]]):
for i in pl.spmd(4): # i = block index
r0 = i * 128
part = pl.matmul(a[r0:r0 + 128, :], b,
out_dtype=pl.FP32)
c = pl.assemble(c, part, [r0, 0])
return c
Runs. Max error vs PyTorch 6.9×10−5.
The orchestration file still submits one task, now with
launch_spec.set_block_num(4). The runtime expands it into four blocks at dispatch, and each
block runs Part I’s kernel on its own 128 rows of A: same K-loop, same ping-pong, same flags. The
dependency dump shows one task, block_num 4, and no edges.
Each block loads all of B. For this shape that is the right trade (B is 128 KiB); for a wide B you would split N as well, and that choice is yours to write.
Rung 2 · more cores, split the reduction
When M and N are too small to feed many cores, split K instead: each core
reduces a slice of K and adds its partial result into c with an atomic add. There are two
obvious ways to write it. They compile, run and give the same answer. Only one of them is parallel.
# A: a loop of device regions (M=N=128, K=2048)
with pl.at(level=CORE_GROUP, name_hint="zero_init"):
c[:] = pl.full([128, 128], dtype=pl.FP32, value=0.0)
for ks in pl.parallel(4):
with pl.at(level=CORE_GROUP, name_hint="split_k"):
k0 = ks * 512
p = pl.matmul(a[:, k0:k0 + 512], b[k0:k0 + 512, :],
out_dtype=pl.FP32)
c = pl.assemble(c, p, [0, 0],
atomic=pl.AtomicType.Add)
# B: one SPMD region (same shapes)
with pl.at(level=CORE_GROUP, name_hint="zero_init"):
c[:] = pl.full([128, 128], dtype=pl.FP32, value=0.0)
for ks in pl.spmd(4):
k0 = ks * 512
p = pl.matmul(a[:, k0:k0 + 512], b[k0:k0 + 512, :],
out_dtype=pl.FP32)
c = pl.assemble(c, p, [0, 0],
atomic=pl.AtomicType.Add)
Serialized. Correct result (8.8×10−5), one core at a time.
zero_init (AIV)
└─wait─▶ split_k[0]
└─wait─▶ split_k[1]
└─wait─▶ split_k[2]
└─wait─▶ split_k[3]
Five tasks, and the runtime chained all of them. Each split_k task declares c as
read-write (the atomic add reads it), so the TensorMap makes each one wait for the previous writer of
c. The hardware atomic in the kernel is real; it just never gets a concurrent writer.
Concurrent. Correct result (9.2×10−5), four cores at once.
zero_init (AIV)
└─wait─▶ split_k one task, 4 blocks
├─ block 0: K[0:512]
├─ block 1: K[512:1024]
├─ block 2: K[1024:1536]
└─ block 3: K[1536:2048]
Two tasks and one edge. The four slices are blocks of a single task, so there is no chain between them to derive: they start together and race on the atomic add in hardware, which is the point.
The rule behind it
simpler orders tasks by the tensors they declare, and an atomic add is still a read-modify-write of the whole output. Tasks that accumulate into the same tensor must therefore be blocks of one SPMD task to run concurrently. Form A is the one a reader coming from GPU programming would write first, and PyPTO’s own matmul tutorial teaches it. The dependency dump is the quickest way to catch this: count the edges.
Rung 3 · bigger than the chip
Make K = 4096. A and B are now 1 MiB each, and Mat holds 512 KiB.
# the Part I program, K = 4096
with pl.at(level=pl.Level.CORE_GROUP):
c[:] = pl.matmul(a, b, out_dtype=pl.FP32)
Compile error at memory allocation (pass 38):
Mat buffer usage (2097152 bytes) exceeds platform limit (524288 bytes)
Pass 12 loads each operand whole into Mat, as in Part I, and nothing between pass 12 and pass 38 tiles at the L1 level. The L0 tiler would happily split K for the cube, but it runs on operands that must already be in Mat.
# block K by hand, 256 at a time
with pl.at(level=pl.Level.CORE_GROUP):
acc = pl.matmul(a[:, 0:256], b[0:256, :],
out_dtype=pl.FP32)
for k in pl.pipeline(1, 16, stage=2):
k0 = k * 256
acc = pl.matmul_acc(acc, a[:, k0:k0 + 256],
b[k0:k0 + 256, :])
c[:] = acc
Runs. Max error 5.2×10−4 (FP16 inputs, K = 4096).
Your stage=2 gives Mat four 64 KiB buffers: two chunks of A and B in
flight, so the next chunk loads from global memory while the current one computes. Inside each chunk
the compiler still tiles for L0 exactly as in Part I.
The compiler tells you what it couldn’t do
Nesting has a cost. Your loop now contains three copies of the compiler’s
own L0 K-loop (a first chunk and two pipeline copies), each asking for its own ping-pong pair, and L0A has
room for two buffers in total. The compile succeeds and prints a performance hint for each:
requested depth 2 … in Left, but only 1 of 2 buffers fit … stages 1 apart share storage
and serialize. The fix is a smaller chunk or a different split, and it is yours to choose.
The other way to say the same thing
pl.pipeline(…, stage=2) is not the only spelling of
ping-pong. pl.MemRef("name", slots=N) declares one allocation with N uniform
slots and picks one per iteration with an ordinary index expression, e.g.
pl.MemRef("ub", slots=2)[i % 2]. Under memory_planner=PTOAS that lowers to a
single ptoas pto.alloc_multi_tile region read through pto.multi_tile_get, and
ptoas derives the per-slot synchronization from the slot index itself. Where
pl.pipeline restructures the loop into a schedule, slots only remove the same-buffer hazard
— the loop stays sequential.
Worth knowing why the compiler’s own matmul pipeline
never uses that form. The slots path requires the loop’s step to be 1 with
a start that is a multiple of N, so that ptoas can match the slot index as a plain
iv % N; the tiled K-loop is emitted with step = k, so it is declined and
replicated instead. And the shape where two slots are live at once is refused at codegen rather than
miscompiled, because the pinned ptoas still mis-synchronizes the prefetch variant of it
(hw-native-sys/PTOAS#1519: fixed in 0.63, regressed in 0.64–0.65). One slot live per
iteration is the shape the region form exists for.
Rung 4 · the layout of the weights
@pl.jit
def matmul_nz(a: pl.Tensor[[128, 512], pl.FP16],
b: pl.Tensor[[512, 128], pl.FP16, pl.NZ],
c: pl.Out[pl.Tensor[[128, 128], pl.FP32]]):
with pl.at(level=pl.Level.CORE_GROUP):
c[:] = pl.matmul(a, b, out_dtype=pl.FP32)
return c
# host side: the bytes must already be in NZ order
def to_nz(x): # [R, C] FP16
r, cols = x.shape
return (x.reshape(r // 16, 16, cols // 16, 16)
.permute(2, 0, 1, 3).contiguous()
.reshape(r, cols))
Runs. Max error 7.6×10−5.
Every load into Mat converts row-major data into the fractal layout on the fly. For weights that are
loaded many times that conversion is repeated work. pl.NZ is not a request to convert: it
is a promise that the bytes in global memory are already in fractal order.
Pass 16 (BlockNzTensorViews, a no-op in Part I) now does something. It rewrites B’s
logical [512, 128] into the blocked [1, 8, 32, 16, 16] the ISA describes
(8 column blocks of 16, 32 row fractals of 16 × 16), and the load of B becomes NZ→NZ
(pto::Layout::NZ in the emitted C++). A stays row-major. Keeping the promise is the
caller’s job; the reorder is the six lines on the left.
Before you go
Static shapes, an output dtype, an output marker and a device region were enough for the stack to produce a tiled, double-buffered, synchronized cube kernel and the host code that launches it.
One pass made a real choice (the L0 tile, from a cost model sized for two buffers). The rest are mechanical: outline, lower to tiles, replicate for ping-pong, reorder loads first, allocate, tag argument directions.
The orchestration library runs on the AI CPU. Edges come from input/output tags through a producer table, and a scope manages buffer lifetime, not ordering.
Two event IDs, one per buffer half, each with a data-ready and a buffer-free direction.
One more line each: pl.spmd for cores, pl.pipeline over your own K-blocks for
L1, pl.NZ for weights. And accumulate across cores inside one SPMD task, or the runtime
will serialize you.
Nothing on this page is privileged
Everything above came from compiling and running the programs on this page with IR dumps and dependency capture turned on, in the PyPTO simulation container, on a laptop.
import os
import torch
import pypto.language as pl
from pypto.ir.pass_manager import PassDumpLevel
from pypto.runtime import RunConfig
# read once per process: set it before compiling
os.environ["PYPTO_PROG_BUILD_DIR"] = "out"
# … the @pl.jit matmul from the top of the page …
a = torch.randn(128, 512, dtype=torch.float16)
b = torch.randn(512, 128, dtype=torch.float16)
c = torch.zeros(128, 128, dtype=torch.float32)
dump = RunConfig(dump_passes=PassDumpLevel.EXPLICIT)
matmul.compile(a, b, c, config=dump)
matmul(a, b, c, config=RunConfig(enable_dep_gen=True))
ref = a.float() @ b.float()
assert torch.allclose(c, ref, rtol=1e-2, atol=1e-2)
Compare against the FP32 product of the FP16 inputs, at an FP16-sized tolerance.
Under out/ you get one directory for the compile and one for the run. The compile directory
holds passes_dump/ (55 files: the parsed input plus one per pass), ptoas/*.pto
and the emitted kernel C++, orchestration/*.cpp, and report/perf_hints.log. The
run directory holds dfx_outputs/deps.json.
cd out/*/passes_dump
prev=; for f in $(ls | sort); do
[ -n "$prev" ] && ! cmp -s "$prev" "$f" && echo "$f"
prev=$f
done
Sixteen names come out for the program on this page. Then read each diff; the table above tells you what to look for.
grep -nE 'set_flag|wait_flag|TEXTRACT|TMATMUL' ptoas/*.cpp gives the instruction stream of the
K-loop figure. For any multi-task program, count the edges in deps.json before
you believe it runs in parallel.
Set K to 256 and pass 19 goes quiet: the operands fit. Set M = N = 256 and it has to split the output too. Set K = 4096 and allocation fails, as in rung 3. The residency explorer above predicts each of these before you compile.