TitanAI: Attention-Free LM Engine on HIP
Attention-free language model training and inference engine written from scratch in HIP for AMD MI300X. Custom BitLinear (1.58-bit) quantized weights, int8 activations, and hand-written dp4a kernels tuned for the 64-thread wavefront architecture.