PyPTO simpler PTOAS one line in, a pipelined cube kernel out every figure from a real compile

One matmul, all the way down

This page starts from the simplest matmul you can write in PyPTO and follows it to the code that runs on an Ascend AI Core. There is no hand-tuned kernel anywhere on the page. Every tile, buffer, loop, and synchronization flag you will see was put there by the compiler stack: PyPTO’s pass pipeline, the simpler runtime, and the PTOAS assembler.

@pl.jit
def matmul(a: pl.Tensor[[128, 512], pl.FP16],
           b: pl.Tensor[[512, 128], pl.FP16],
           c: pl.Out[pl.Tensor[[128, 128], pl.FP32]]):
    with pl.at(level=pl.Level.CORE_GROUP):
        c[:] = pl.matmul(a, b, out_dtype=pl.FP32)
    return c
16 of 54compiler passes change this program; the other 38 have nothing to do, and the page says why for each
2 × 2ping-pong buffers in the cube’s operand ports, allocated without being asked for
4K-iterations of 128, chosen by a cost model because A (128 KiB) cannot fit the 64 KiB operand buffer
2 programscome out: a cube kernel for the AI Core and an orchestration library the runtime loads on the AI CPU

Part I follows that line down, stage by stage, and ends at the instruction stream. Part II asks the question an HPC reader asks next: this kernel runs on one core. Which component makes it use all of them, keep weights in the hardware’s native layout, or handle matrices that don’t fit on the chip? For each of those, the answer is a specific pass, a specific runtime feature, or “you, in one more line of DSL”. Provenance. Structure is observed: every pass diff, buffer address, instruction, synchronization flag and dependency edge below is copied from one compile of this program and from the runtime’s own dependency dump. Target a2a3sim, backend Ascend 910B, PyPTO ee49fcea. The simulator runs the real kernel (the result matches PyTorch to 7.6×10−5) but is not timing-accurate, so nothing on this page is a cycle count, a bandwidth, or a speedup.

Before the compiler runs

What the one line already says

The program names no memory space, no tile size and no loop. It still commits to four things, and each one decides something later on the page.

pl.Tensor[[128, 512], pl.FP16]
Shapes are static and part of the type. Every decision below (whether the matmul fits the chip, how to tile it, where each buffer goes) is made at compile time from these numbers. Change a shape and you get a different kernel.
out_dtype=pl.FP32
The cube multiplies FP16 operands and accumulates in FP32. There is no cast op anywhere in the compiled program: the FP32 result is simply the type of the accumulator tile. Asking for an FP16 output would throw away precision the hardware already paid for.
pl.Out[…] and return c
c is a caller-owned buffer the program writes. That marking survives all the way to the runtime, which uses it to decide that c needs no copy to the device before the run and a, b need no copy back after it.
pl.at(level=CORE_GROUP)
“This region is device work.” Everything inside becomes one task for one core group; the function around it becomes host-side orchestration. That split is the first real thing the compiler does (pass 10), and it is why two programs come out.

Why placement must be explicit on this target

For this page’s Ascend 910B target (a 220x-class separation-mode device), the relevant execution model is MPMD-like: the AIC (Cube) and AIV (Vector) cores have independent scalar control paths and can load separate instruction streams. Data exchanged between them follows the documented global-memory path on this target, while explicit cross-core flags provide ordering. PyPTO therefore cannot rely on a transparent GPU-style cache hierarchy for tensor tiles; placement and movement have to be exposed or inferred. A manual six-statement matmul makes that contract visible: two GM→Mat loads, two Mat→Left/Right moves, one cube multiply producing Acc, and one store back to GM. The whole-tensor example on this page is a separate route: the AutoTileMatmulL0 pass chooses legal L0 tiles with a cost model, then PyPTO emits PTO MLIR and PTOAS lowers it to C++ calls into PTO-ISA before compilation.

The map

Three tools, two programs, one run

PyPTO’s passes turn the program into two functions. The device kernel goes through PTOAS to C++ calls into PTO-ISA for the AI Core. The orchestration function becomes C++ against simpler’s API and runs on the AI CPU. simpler brings the two back together at run time. Step through once; each later section expands one box.

STAGE 0 / 7
kernel path — AI Core orchestration path — AI CPU 1pl.matmulDSL 254 passesPyPTO IR 3.ptoPTOAS MLIR 4kernel.cpp+ sync flags 5orchestration.cppcalls simpler’s task API 6compile.o + .so 7simplerrun → cube
The fork after stage 2 is the two-program split made concrete: the compiler emits a device program and a host-side program, and the runtime joins them.
Stage 2 · the territory

What the passes are placing tiles into

The memory spaces, and who can reach them

Before watching the compiler place tiles, look at where it can put them. Each on-chip memory space is its own SRAM with its own capacity, its own owning core, and a short list of legal ways in and out. They are separate buffers, not levels of a cache: L0A is not “inside” L1 the way a CPU’s L1 sits inside its L2. Click or tab to any space or arrow.

this kernel’s path legal, unused here rejected by the compiler
Stage 2 · PyPTO passes

The whole pipeline, one line per pass

Sixteen passes out of fifty-four

PyPTO lowers the program through 54 passes over one IR. Tensor ops and tile ops live in the same tree; pass 12 is where one turns into the other. Diffing every pass’s output against its input shows that 16 of them change this program. The table lists all 54: what each pass is for in general, and what it did here or why it had nothing to do. Highlighted rows changed the IR; click one to open its diff in the stepper below.

#PassWhat it doesThis programDiff

Phases are this page’s grouping of the pass manager’s order. Pass numbers are the pass manager’s invocation order, which is what passes_dump/NN_after_<Pass>.py files are named after; the pass documentation numbers each one lower (this page’s pass 19 is doc 18, pass 37 is doc 36). The pass list, order and descriptions come from the compiler’s own pass registry and documentation; the “this program” column comes from diffing consecutive IR dumps of this compile.

PASS 0 / 16

Diffs are the real IR text of consecutive dumps, shown as unified diffs with three lines of context. For readability the dumps were run through ruff format (88 columns) before diffing, tile layout annotations (pl.TileView(…)) are dropped, and pl.const(N, pl.INT64) is printed as N. The +/− counts are of the displayed lines.

Who wrote Left and Right?

A pass called InferTileMemorySpace exists, and you would expect it to place the operands. Here it doesn’t. Pass 12 already put A and B in Mat and the result in Acc, and pass 19 wrote target_memory=Left/Right explicitly on its extracts. It has to: it runs two passes before inference, at a point where every tile is 2D but no space has been inferred yet. By pass 21 there is nothing left to infer.

Double buffering is three passes, not one

AutoTileMatmulL0 (19) only asks for two stages. LowerPipelineLoops (32) copies the loop body so consecutive iterations touch different buffers. CanonicalizeIOOrder (33) moves the second copy’s loads above the first copy’s matmul, so the DMA engine has work queued while the cube computes. Then MemoryReuse (37) is told not to merge the two copies’ buffers, which is the only reason two of each survive to allocation.

Stage 2 · pass 19

The one real decision in the pipeline

Why the compiler picked 128 × 128 × 128

Of the sixteen passes that changed something, one made a choice with alternatives: AutoTileMatmulL0 picked the L0 tile. The operands as written cannot sit in the cube’s operand buffers at all. A is 128 × 512 FP16, which is 128 KiB, and L0A holds 64 KiB. Something has to split K.

m128
n128
k512
Operands

FP16 operands (2 bytes), FP32 accumulator (4 bytes), Ascend 910B capacities. This is residency arithmetic only. It shows which shapes are legal, not which legal shape is fastest.

The rule that picks k = 128

When both operands stream through the cube (the output-stationary case, where the accumulator stays put and A and B slices flow past it), the chooser plans for two copies of each operand tile from the start. It sizes k against half of L0A: 64 KiB / (2 bytes × 2 stages) = 16 384 elements, so with m = 128 the largest legal k is 128. k = 256 would fit one buffer exactly and leave no room for the second. The pick fills L0A and L0B to exactly 100 %.

A search with a cost model, not a table

The chooser enumerates every legal 16-aligned (m, n, k), crossed with which operand stays resident and whether the accumulator is double-buffered, and scores each candidate with a roofline model of cube cycles, operand-load cycles (weighted by the unequal L0A and L0B port bandwidths) and accumulator drain cycles. It returns the minimum. The model is checked against a brute-force re-enumeration of itself, which proves the search is correct. Nobody has timed the pick against alternatives on hardware.

16 is the hardware’s unit, not 128

The chosen 128 × 128 tile is not a hardware parameter. The cube works on fractal boxes, and inside every on-chip buffer a tile is stored as a grid of them: a 512-byte box (16 rows of 32 bytes) for the FP16 operand tiles, a 1024-byte box for the FP32 accumulator (visible as fractal=512 and fractal=1024 in the compiled tile types). The tiler’s constraints are stated in those units: every tile dimension a multiple of 16, and K 16-aligned or the pass declines. 128 is just the largest multiple of 16 that satisfies the double-buffered budget above.

The chooser splits K before it ever splits M or N. Splitting K costs nothing extra: the partial sums accumulate in one L0C tile and drain to memory once. Splitting M or N means one drain per output sub-tile. So it tiles the output only when the output itself doesn’t fit L0C, and here a 128 × 128 FP32 accumulator is 64 KiB of a 128 KiB buffer, so it doesn’t.

Stage 2 · passes 35–38

The territory, as the allocator filled it

Seven buffers and their addresses

Each panel is one memory space, drawn to that space’s real capacity. Time runs downward through the kernel’s statements, so a block’s height is how long its buffer is live. Every space starts its own address line at zero, which is what “separate buffers, not a hierarchy” means in practice. Hover a block or a statement.

Full, on purpose

L0A and L0B are at 100 % and L0C at 50 %. That is the tile choice from the previous section, showing up as addresses: ping at offset 0, pong at 32 KiB, nothing left over. Mat holds all of A and B (256 KiB of 512) for the whole kernel, because the program fits on the chip. Part II shows what happens when it doesn’t.

Stages 5 & 7 · simpler

The runtime half of the program

What simpler does with the other program

Pass 10 split the program in two, and so far the page has followed the kernel. The other half, the orchestration function, is also compiled. It becomes a small C++ library that runs on the device’s AI CPU, not on the host, and its only job is to tell simpler’s scheduler what tasks exist and which tensors each one reads and writes. This section follows that half from the compiler to the moment the cube starts.

Four passes build the seam

Nothing about the runtime is decided in a separate tool. Four of the sixteen passes that changed the IR exist to shape the orchestration side, and each one maps to one line of the generated file below.

PassIR change hereWhat it becomes in the orchestration file
10 OutlineIncoreScopesthe pl.at body becomes kernel matmul_matmul; the caller keeps one callone task, params_t0, and one submit call
25 ExpandMixedKernelkernel retyped InCore → AIC (pure cube, no vector half)rt_submit_aic_task, not the mixed-cluster form
42 DeriveCallDirectionsarg_directions = [input, input, output_existing]add_input, add_input, add_output
51 MaterializeRuntimeScopesbody wrapped in one pl.scope()one SIMPLER_SCOPE() { … }

The generated orchestration file

This is the whole file, as emitted. Click a line.

SIMPLER_SCOPE is not a barrier

It reads like a fork-join block, and it isn’t one. Closing a scope doesn’t wait for the scope’s tasks to finish. It drops the scope’s reference on each task it submitted, which is what allows the runtime to recycle those tasks’ ring slots and intermediate buffers once their consumers are done. Ordering between tasks comes from somewhere else entirely: the tensors each task reads and writes, as the next part shows.

Stage 7 · one run

From the Python call to the cube and back

Three programs, one run

A run involves three separately compiled programs: simpler’s host library, the orchestration library on the AI CPU, and the kernel on the AI Core. Step through one call of matmul(a, b, c).

STEP 0 / 9
Each box is one step; arrows are hand-offs between programs. On a2a3sim the same three programs run as host threads, with the AI Core’s control registers replaced by plain memory.

The dependency graph simpler recorded

With dependency capture on, the runtime writes down every task it was given and every edge it derived. For this program: one task on the cube, three tensor arguments, no edges. The direction tags are the ones pass 42 derived.

How edges get made

There is no edge list in the orchestration code. When a task is submitted, the runtime looks up each of its input tensors in a producer table (the TensorMap) and adds an edge from whichever earlier task last wrote it, then records this task as the producer of its outputs. A task with no producers goes straight onto the ready queue for its core type. So the direction tags from pass 42 aren’t metadata; they are the dependency graph.

That has consequences once there is more than one task. Part II has a case where a documented multi-core pattern quietly runs on one core at a time, and the reason is visible in exactly this kind of dump.

Stages 3 & 4 · PTOAS

From tile IR to C++, and where the synchronization comes from

PTOAS adds the one thing PyPTO left out

After pass 54 the kernel is printed as an MLIR module in the PTO dialect and handed to PTOAS, the assembler. In that module the memory spaces are still part of each tile’s type (loc=mat, loc=left, loc=acc), and so are the layouts from the memory-space table. What the module does not contain is a single synchronization instruction. PyPTO has no intra-core sync pass. PTOAS inserts every set_flag and wait_flag the kernel needs, and it is the only component that does.

in — the loop body of the .pto module (types shortened)

scf.for %ko = 0 to 512 step 256 {
  %a0 = pto.alloc_tile addr = 0     : tile_buf<loc=left, f16, 128x128>
  pto.textract ins(%a_mat, 0, %ko) outs(%a0)
  %b0 = pto.alloc_tile addr = 0     : tile_buf<loc=right, f16, 128x128>
  pto.textract ins(%b_mat, %ko, 0) outs(%b0)
  %a1 = pto.alloc_tile addr = 32768 : tile_buf<loc=left, …>
  pto.textract ins(%a_mat, 0, %ko+128) outs(%a1)
  %b1 = pto.alloc_tile addr = 32768 : tile_buf<loc=right, …>
  pto.textract ins(%b_mat, %ko+128, 0) outs(%b1)
  scf.if %ko == 0 { pto.tmatmul(%a0, %b0) }
         else { pto.tmatmul.acc(%acc, %a0, %b0) }
  scf.if %ko == -128 { pto.tmatmul(%a1, %b1) }
         else { pto.tmatmul.acc(%acc, %a1, %b1) }
}
// no set_flag / wait_flag anywhere in the module

out — the same loop in the emitted C++ (declarations dropped)

wait_flag(PIPE_MTE2, PIPE_MTE1, EVENT_ID0);  // A, B in Mat
for (ko = 0; ko < 512; ko += 256) {
  wait_flag(PIPE_M, PIPE_MTE1, EVENT_ID0);   // ping free?
  pipe_barrier(PIPE_MTE1);
  TEXTRACT(a0, a_mat, 0, ko);
  TEXTRACT(b0, b_mat, ko, 0);
  set_flag(PIPE_MTE1, PIPE_M, EVENT_ID0);    // ping full
  wait_flag(PIPE_M, PIPE_MTE1, EVENT_ID1);   // pong free?
  TEXTRACT(a1, a_mat, 0, ko + 128);
  TEXTRACT(b1, b_mat, ko + 128, 0);
  set_flag(PIPE_MTE1, PIPE_M, EVENT_ID1);    // pong full
  wait_flag(PIPE_MTE1, PIPE_M, EVENT_ID0);
  if (ko == 0) TMATMUL(acc, a0, b0);
  else         TMATMUL_ACC(acc, acc, a0, b0);
  set_flag(PIPE_M, PIPE_MTE1, EVENT_ID0);    // ping free
  wait_flag(PIPE_MTE1, PIPE_M, EVENT_ID1);
  if (ko == -128) TMATMUL(acc, a1, b1);
  else            TMATMUL_ACC(acc, acc, a1, b1);
  set_flag(PIPE_M, PIPE_MTE1, EVENT_ID1);    // pong free
}

Both listings are excerpts of the real files from this compile, with long template types and variable declarations removed and SSA names shortened (%t__tile_l0_a → %a0). The comments on the right are this page’s.

Read the event IDs as buffer names

EVENT_ID0 is the ping buffer pair and EVENT_ID1 the pong pair. Each pair has two flags going in opposite directions. MTE1→M says “this buffer is full, compute may read it” (read-after-write). M→MTE1 says “compute has finished with this buffer, the DMA may overwrite it” (write-after-read). Before the loop, PTOAS pre-sets both M→MTE1 flags, so the first two extracts don’t wait for a matmul that never happened; after the loop it drains them.

How PTOAS gets there

The sync pass runs six stages in order. It translates the module into a list of operations tagged with the pipe each runs on and the buffers it touches, finds every cross-pipe hazard, moves sync points to cheaper positions, removes syncs that others already cover, assigns hardware event IDs, and emits the calls. The loop bound and both ko == … tests come through unchanged; the second one can never be true, because pass 49 left it unfolded.

InstructionPipeHardwareIssued
TLOAD × 2MTE2DMA: global memory → L1once, before the loop
TEXTRACT × 4 per tripMTE1DMA: L1 → L0A / L0B8 in total
TMATMUL / TMATMUL_ACCMthe cube1 + 3 in total
TSTOREFIXFixPipe: L0C → global memoryonce, after the loop
Stage 7 · on the core

What stage=2 actually buys

Four iterations on four pipes

Each pipe (the two DMA engines, the cube, the FixPipe) executes its own instructions in order, and the pipes run in parallel with each other. The only thing that holds one pipe back for another is a flag. This figure places every instruction of the kernel at the earliest step its flags allow, assuming every instruction takes one step. Switch to stage=1 to see the same kernel with a single buffer pair.

Operand buffers Step
read-after-write: data ready write-after-read: buffer free again accumulator chain ping (offset 0, EVENT_ID0) pong (offset 32 KiB, EVENT_ID1)

What the figure shows

With two buffer pairs, the extract for iteration i+1 goes into the other half while matmul i is still reading the first, so MTE1 and the cube stay busy on the same steps. The dashed violet edges are the new constraint double buffering introduces: iteration i+2 reuses iteration i’s half, so its extract waits for matmul i to finish. The amber chain is the one dependency no buffering can remove: all four matmuls accumulate into the same L0C tile, in order.

What it does not show

Positions are dependency steps, not time. On real hardware a 128 × 128 × 128 FP16 matmul and a 32 KiB L1→L0 copy take different numbers of cycles, and whether MTE1 or the cube is the bottleneck depends on those numbers. The simulator this page used runs the kernel correctly but does not model timing. The step counts are a statement about the kernel’s dependency structure, which is exactly what the flags encode.


Part II

The extra mile, and who walks it

Part I’s kernel is correct, tiled for L0, double-buffered and synchronized, and it runs on one of the 910B’s 24 cube cores with operands that fit on the chip. Everything an HPC matmul adds beyond that has an owner. Here is who owns what, checked against the compiler, the runtime and the assembler source, and then each missing piece added to the same program and compiled.

OptimizationWho does itHowIn Part I?
L0 tiling: K-loop, and M/N split when the output overflows L0CPyPTO pass 19automatic cost-model searchyes: k = 128
L0A/L0B ping-pong inside that loopPyPTO passes 19, 32, 33, 37automatic once the matmul is tiledyes: 2 × 2 buffers
Intra-core synchronization and event IDsPTOAS sync passautomatic always enabled by PyPTOyes: 2 events
Fractal layout for tiles in Mat / L0PyPTO tile conversionautomatic converted during the loadyes
Accumulator (L0C) double bufferingPyPTO passes 19, 32, 33, 37automatic or opt-in automatic under memory_planner=PTOAS and DSA_RP; a marked-experimental flag under the default PyPTO plannerno — k = 128 is split-K, and dbC = 2 requires the full-K emitter
Using more than one core: M/N split or split-Kyou, in the DSL; simpler dispatchesone more line pl.spmd(n)no — rungs 1, 2
L1 tiling: operands bigger than the 512 KiB Matnobodycompile error you block K yourselfno — rung 3
Double buffering of your own loops (GM→L1)you; PyPTO lowers itone more line pl.pipeline(…, stage=2)no — rung 3
Weights stored in the hardware layoutyou; PyPTO pass 16 plans for ita type annotation pl.NZno — rung 4
Epilogue fusion (bias, ReLU, quantize on writeback)you, at tile leveltile-level store optionsno
Overlap between taskssimplerautomatic for tasks you created, ordered by their tensorsone task: nothing to overlap

What those four passes each do

The ping-pong row names four passes because no single pass owns this optimization. Reading them in order is the whole mechanism — and the fourth one is the one that people usually assume is the second.

  • 19 AutoTileMatmulL0 requests it. The tiled K-loop is emitted as ForKind::Pipeline with pipeline_stages=2. That 2 is the operand depth — L0A and L0B — not the accumulator. It is also why k = 128 and not 512: the chooser sizes the tile against half of each operand buffer, which is the arithmetic in the previous section.
  • 32 LowerPipelineLoops replicates the body. Two copies per outer iteration, each with fresh def-vars, and each clone's tile-producing call tagged with a pipeline_membership (group, stage) pair.
  • 33 CanonicalizeIOOrder shapes the schedule — scalars first, then the loads clustered ahead of compute, so the load engine runs ahead of the cube. It only reorders statements; it does not create the buffer separation.
  • 37 MemoryReuse is where the separation is actually enforced. The two clones are sequential in program order, so their tiles have disjoint lifetimes — exactly the condition the reuse packer coalesces on. Left alone it would fuse the ping-pong back into a single buffer. The pipeline_membership tag is what forbids that, and it does so only as deep as the memory space can afford.

The accumulator is deliberately left untagged: it is written by the serialized cube, one tile's MAD retiring before the next begins, so a single L0C buffer is already correct. That is why this page shows one Acc @ 0 rather than a pair.

The pattern

The compiler automates the single-core, fits-on-chip story completely. Everything about many cores and data bigger than the chip is written by the programmer today, as small DSL additions to the same program. That split is deliberate: the assembler’s design notes rule out inferring multi-buffering on its own, the accumulator double buffer is experimental under the default planner (it is automatic under memory_planner=PTOAS and DSA_RP, and needs the full-K emitter either way), and nothing in the compiler splits one device region across cores. Production kernels built on PyPTO write all four rungs below by hand.

The switch this page did not have to touch

Everything above was compiled with the default memory_planner=PYPTO. That setting decides who owns on-chip memory, and it is the one knob that can also move the tile the chooser picks.

PYPTO (default)PTOAS
Addressesbaked addr on every tile; ptoas runs at --pto-level=level3addr-less tiles, placed by ptoas at --pto-level=level2
Reuse and placementPyPTO MemoryReuse + AllocateMemoryAddrPTOAS PlanMemory, with both of those skipped
L0C double bufferopt-in flagautomatic, if the chooser picks it

It can change the tile. dbC = 2 requires the full-K emitter — l0_tile_chooser.cpp sets require_full_k = !is_os || r.dbc — and halves the L0C accumulator budget (l0c_bytes / (bytes_c × 2)), so making it available re-ranks the candidate shapes. On this program it changed nothing: the same 128 × 128, k = 128, one accumulator, at identical offsets, under all three configurations. That is the useful part — Part I is a default-planner picture, and it holds because k = 128 split-K won on modelled cost, not because a full-K dbC = 2 candidate was illegal.

Rung 1 · more cores, split the output

Four blocks with pl.spmd

# M = 512, N = 128, K = 512
@pl.jit
def matmul_spmd(a: pl.Tensor[[512, 512], pl.FP16],
                b: pl.Tensor[[512, 128], pl.FP16],
                c: pl.Out[pl.Tensor[[512, 128], pl.FP32]]):
    for i in pl.spmd(4):          # i = block index
        r0 = i * 128
        part = pl.matmul(a[r0:r0 + 128, :], b,
                         out_dtype=pl.FP32)
        c = pl.assemble(c, part, [r0, 0])
    return c

Runs. Max error vs PyTorch 6.9×10−5.

The orchestration file still submits one task, now with launch_spec.set_block_num(4). The runtime expands it into four blocks at dispatch, and each block runs Part I’s kernel on its own 128 rows of A: same K-loop, same ping-pong, same flags. The dependency dump shows one task, block_num 4, and no edges.

Each block loads all of B. For this shape that is the right trade (B is 128 KiB); for a wide B you would split N as well, and that choice is yours to write.

Rung 2 · more cores, split the reduction

Split-K, and the version that runs on one core at a time

When M and N are too small to feed many cores, split K instead: each core reduces a slice of K and adds its partial result into c with an atomic add. There are two obvious ways to write it. They compile, run and give the same answer. Only one of them is parallel.

# A: a loop of device regions  (M=N=128, K=2048)
with pl.at(level=CORE_GROUP, name_hint="zero_init"):
    c[:] = pl.full([128, 128], dtype=pl.FP32, value=0.0)
for ks in pl.parallel(4):
    with pl.at(level=CORE_GROUP, name_hint="split_k"):
        k0 = ks * 512
        p = pl.matmul(a[:, k0:k0 + 512], b[k0:k0 + 512, :],
                      out_dtype=pl.FP32)
        c = pl.assemble(c, p, [0, 0],
                        atomic=pl.AtomicType.Add)
# B: one SPMD region  (same shapes)
with pl.at(level=CORE_GROUP, name_hint="zero_init"):
    c[:] = pl.full([128, 128], dtype=pl.FP32, value=0.0)
for ks in pl.spmd(4):
    k0 = ks * 512
    p = pl.matmul(a[:, k0:k0 + 512], b[k0:k0 + 512, :],
                  out_dtype=pl.FP32)
    c = pl.assemble(c, p, [0, 0],
                    atomic=pl.AtomicType.Add)

Serialized. Correct result (8.8×10−5), one core at a time.

zero_init (AIV)
 └─wait─▶ split_k[0]
           └─wait─▶ split_k[1]
                     └─wait─▶ split_k[2]
                               └─wait─▶ split_k[3]

Five tasks, and the runtime chained all of them. Each split_k task declares c as read-write (the atomic add reads it), so the TensorMap makes each one wait for the previous writer of c. The hardware atomic in the kernel is real; it just never gets a concurrent writer.

Concurrent. Correct result (9.2×10−5), four cores at once.

zero_init (AIV)
 └─wait─▶ split_k   one task, 4 blocks
            ├─ block 0: K[0:512]
            ├─ block 1: K[512:1024]
            ├─ block 2: K[1024:1536]
            └─ block 3: K[1536:2048]

Two tasks and one edge. The four slices are blocks of a single task, so there is no chain between them to derive: they start together and race on the atomic add in hardware, which is the point.

The rule behind it

simpler orders tasks by the tensors they declare, and an atomic add is still a read-modify-write of the whole output. Tasks that accumulate into the same tensor must therefore be blocks of one SPMD task to run concurrently. Form A is the one a reader coming from GPU programming would write first, and PyPTO’s own matmul tutorial teaches it. The dependency dump is the quickest way to catch this: count the edges.

Rung 3 · bigger than the chip

When the operands don’t fit in L1

Make K = 4096. A and B are now 1 MiB each, and Mat holds 512 KiB.

# the Part I program, K = 4096
with pl.at(level=pl.Level.CORE_GROUP):
    c[:] = pl.matmul(a, b, out_dtype=pl.FP32)

Compile error at memory allocation (pass 38):
Mat buffer usage (2097152 bytes) exceeds platform limit (524288 bytes)

Pass 12 loads each operand whole into Mat, as in Part I, and nothing between pass 12 and pass 38 tiles at the L1 level. The L0 tiler would happily split K for the cube, but it runs on operands that must already be in Mat.

# block K by hand, 256 at a time
with pl.at(level=pl.Level.CORE_GROUP):
    acc = pl.matmul(a[:, 0:256], b[0:256, :],
                    out_dtype=pl.FP32)
    for k in pl.pipeline(1, 16, stage=2):
        k0 = k * 256
        acc = pl.matmul_acc(acc, a[:, k0:k0 + 256],
                            b[k0:k0 + 256, :])
    c[:] = acc

Runs. Max error 5.2×10−4 (FP16 inputs, K = 4096).

Your stage=2 gives Mat four 64 KiB buffers: two chunks of A and B in flight, so the next chunk loads from global memory while the current one computes. Inside each chunk the compiler still tiles for L0 exactly as in Part I.

The compiler tells you what it couldn’t do

Nesting has a cost. Your loop now contains three copies of the compiler’s own L0 K-loop (a first chunk and two pipeline copies), each asking for its own ping-pong pair, and L0A has room for two buffers in total. The compile succeeds and prints a performance hint for each: requested depth 2 … in Left, but only 1 of 2 buffers fit … stages 1 apart share storage and serialize. The fix is a smaller chunk or a different split, and it is yours to choose.

The other way to say the same thing

pl.pipeline(…, stage=2) is not the only spelling of ping-pong. pl.MemRef("name", slots=N) declares one allocation with N uniform slots and picks one per iteration with an ordinary index expression, e.g. pl.MemRef("ub", slots=2)[i % 2]. Under memory_planner=PTOAS that lowers to a single ptoas pto.alloc_multi_tile region read through pto.multi_tile_get, and ptoas derives the per-slot synchronization from the slot index itself. Where pl.pipeline restructures the loop into a schedule, slots only remove the same-buffer hazard — the loop stays sequential.

Worth knowing why the compiler’s own matmul pipeline never uses that form. The slots path requires the loop’s step to be 1 with a start that is a multiple of N, so that ptoas can match the slot index as a plain iv % N; the tiled K-loop is emitted with step = k, so it is declined and replicated instead. And the shape where two slots are live at once is refused at codegen rather than miscompiled, because the pinned ptoas still mis-synchronizes the prefetch variant of it (hw-native-sys/PTOAS#1519: fixed in 0.63, regressed in 0.64–0.65). One slot live per iteration is the shape the region form exists for.

Rung 4 · the layout of the weights

Store B the way the cube reads it

@pl.jit
def matmul_nz(a: pl.Tensor[[128, 512], pl.FP16],
              b: pl.Tensor[[512, 128], pl.FP16, pl.NZ],
              c: pl.Out[pl.Tensor[[128, 128], pl.FP32]]):
    with pl.at(level=pl.Level.CORE_GROUP):
        c[:] = pl.matmul(a, b, out_dtype=pl.FP32)
    return c

# host side: the bytes must already be in NZ order
def to_nz(x):                    # [R, C] FP16
    r, cols = x.shape
    return (x.reshape(r // 16, 16, cols // 16, 16)
             .permute(2, 0, 1, 3).contiguous()
             .reshape(r, cols))

Runs. Max error 7.6×10−5.

Every load into Mat converts row-major data into the fractal layout on the fly. For weights that are loaded many times that conversion is repeated work. pl.NZ is not a request to convert: it is a promise that the bytes in global memory are already in fractal order.

Pass 16 (BlockNzTensorViews, a no-op in Part I) now does something. It rewrites B’s logical [512, 128] into the blocked [1, 8, 32, 16, 16] the ISA describes (8 column blocks of 16, 32 row fractals of 16 × 16), and the load of B becomes NZ→NZ (pto::Layout::NZ in the emitted C++). A stays row-major. Keeping the promise is the caller’s job; the reorder is the six lines on the left.

Before you go

What you have seen

one line

Everything below it was derived

Static shapes, an output dtype, an output marker and a device region were enough for the stack to produce a tiled, double-buffered, synchronized cube kernel and the host code that launches it.

PyPTO

16 of 54 passes did the work

One pass made a real choice (the L0 tile, from a cost model sized for two buffers). The rest are mechanical: outline, lower to tiles, replicate for ping-pong, reorder loads first, allocate, tag argument directions.

simpler

The dependency graph is the argument tags

The orchestration library runs on the AI CPU. Edges come from input/output tags through a producer table, and a scope manages buffer lifetime, not ordering.

PTOAS

Every flag is inserted by the assembler

Two event IDs, one per buffer half, each with a data-ready and a buffer-free direction.

the extra mile

Many cores and big data are yours to write

One more line each: pl.spmd for cores, pl.pipeline over your own K-blocks for L1, pl.NZ for weights. And accumulate across cores inside one SPMD task, or the runtime will serialize you.

Nothing on this page is privileged

Do this yourself

Everything above came from compiling and running the programs on this page with IR dumps and dependency capture turned on, in the PyPTO simulation container, on a laptop.

Compile with per-pass dumps, run with dependency capture

import os

import torch
import pypto.language as pl
from pypto.ir.pass_manager import PassDumpLevel
from pypto.runtime import RunConfig

# read once per process: set it before compiling
os.environ["PYPTO_PROG_BUILD_DIR"] = "out"

# … the @pl.jit matmul from the top of the page …

a = torch.randn(128, 512, dtype=torch.float16)
b = torch.randn(512, 128, dtype=torch.float16)
c = torch.zeros(128, 128, dtype=torch.float32)

dump = RunConfig(dump_passes=PassDumpLevel.EXPLICIT)
matmul.compile(a, b, c, config=dump)
matmul(a, b, c, config=RunConfig(enable_dep_gen=True))

ref = a.float() @ b.float()
assert torch.allclose(c, ref, rtol=1e-2, atol=1e-2)

Compare against the FP32 product of the FP16 inputs, at an FP16-sized tolerance.

Know what landed where

Under out/ you get one directory for the compile and one for the run. The compile directory holds passes_dump/ (55 files: the parsed input plus one per pass), ptoas/*.pto and the emitted kernel C++, orchestration/*.cpp, and report/perf_hints.log. The run directory holds dfx_outputs/deps.json.

Find the passes that matter

cd out/*/passes_dump
prev=; for f in $(ls | sort); do
  [ -n "$prev" ] && ! cmp -s "$prev" "$f" && echo "$f"
  prev=$f
done

Sixteen names come out for the program on this page. Then read each diff; the table above tells you what to look for.

Read the flags, then count the edges

grep -nE 'set_flag|wait_flag|TEXTRACT|TMATMUL' ptoas/*.cpp gives the instruction stream of the K-loop figure. For any multi-task program, count the edges in deps.json before you believe it runs in parallel.

Push the shape and watch the compiler change its mind

Set K to 256 and pass 19 goes quiet: the operands fit. Set M = N = 256 and it has to split the output too. Set K = 4096 and allocation fails, as in rung 3. The residency explorer above predicts each of these before you compile.