Skip to content
View titoatwork's full-sized avatar

Highlights

  • Pro

Block or report titoatwork

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
titoatwork/README.md

Fourth-year CS undergrad. I work on making models cheaper to run and measurements harder to fake.

Mostly GPU kernels, low-precision inference, and evaluation harnesses that don't flatter their author.

Things that exist

slipstream · Python / Triton
LLM inference engine, built from scratch. The contribution sits at the scheduling layer — a memory-aware, output-length-predictive policy under hard KV-cache limits, where production engines still schedule blind. Companion to qgemm-mx: same bandwidth wall, one layer up.

qgemm-mx · C++ / CUDA
Block-scaled FP4 on GPUs with no native FP4. MXFP4 moves 1.88x fewer bytes per weight than FP8 and currently delivers roughly none of the speedup. Measuring where it goes.

lfx-firstanalysis · Python
Every published figure re-derives offline from a committed artifact, or the build fails. ./verify.sh, two dependencies, no API key.

COLIDE · Python / CUDA
CNN-BiLSTM intrusion detection over network flow, with the inference path moved to GPU.

riscv/riscv-unified-db · Ruby / YAML
18 patches merged, none rejected. Most of them closed a defect my own measurement had surfaced.

Earlier

Transformer inference on DGX H100 clusters. A fused Triton softmax at 3.53x PyTorch, built with A. Agrawal. A hedging agent that beats Black-Scholes mainly by trading less.

Currently reading about

Ternary quantization, wave quantization on small batch sizes, and why nobody agrees on what a benchmark measured.

Pinned Loading

  1. COLIDE COLIDE Public

    COLIDE: CUDA-optimized CNN-BiLSTM IoT IDS (Option A). Principal BoT F1 0.9780±0.0033. Manuscript-ready evidence pack; see docs/CHERAN_MANUSCRIPT_HANDOFF.md

    Python 1

  2. qgemm-mx qgemm-mx Public

    Recovering block-scaled FP4 (MXFP4/NVFP4) throughput on GPUs without native FP4 support. Measurement study + Hopper kernels.

    Cuda 1

  3. slipstream slipstream Public

    High-throughput LLM inference engine built from scratch, with a novel memory-aware scheduler (Horizon).

    Python 1

  4. iemAnshuman/Triton_1.58 iemAnshuman/Triton_1.58 Public

    Python