The Spark codecs are often bandwidth-bound. The codec usually reads an uncompressed RGBA8 texture and outputs a block-compressed texture that, depending on the format, is 2–8 times smaller. Reading the input is what dominates bandwidth. When bandwidth-bound, execution stalls waiting for input data, while the GPU tries to hide that latency by overlapping loads for some blocks with computation on others.
Mobile GPUs render one tile at a time and keep its data in on-chip memory to save bandwidth and power. Wouldn’t it be great if our codecs could also use this tile memory as their input? That’s exactly what Metal tile shaders allow us to do.
Continue reading →
