Data platforms. AI systems. The infrastructure between them.
delta at Google Cloud, leading Data & AI innovation and transformation across Southeast Asia. I publish research on multi-agent systems, inference optimization, and AI safety, and write the practitioner's notes behind it, from inside real production systems.
01 Now
ICLR 2026 paper on LLM safety, and compute primitives for orbital environments.
New ventures at the frontier; conversations with research labs and founders.
Fused MoE kernels, circuit tracing in production, and bets on model honesty.
02 Latest
Beating FP16 with 4-bit Weights: A Portable W4A16 GEMM in Triton
I wrote a 4-bit weight-only GEMM in pure Triton. The fast W4A16 kernels are all CUDA, so this one runs on NVIDIA and AMD. It beats cuBLAS FP16 by 1.1 to 1.3x in the decode regime, and the road there was mostly me bein...
Read the essay →03 Selected writing
04 Research focus
Multi-agent systems
Coordination, debate, and verification for AI that runs unattended.
/ 02Inference optimization
Fused MoE dispatch, speculative decoding, custom Triton kernels.
/ 03AI safety & interpretability
Circuit tracing, activation steering, sandbagging detection.
/ 04Distributed systems
Formal synthesis and correct-by-construction architectures.
05 Instruments
All instruments →06 Selected publications