Skip to content
View SkyCWO's full-sized avatar
🏠
Working from home
🏠
Working from home
  • @SJTU-AI-Lab
  • Shanghai, China
  • 12:47 (UTC +09:00)

Block or report SkyCWO

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
SkyCWO/README.md

Haoyu Zhang

Graduate student at SJTU, working at the intersection of speech, vision, and language. I'm mostly interested in how to make large multimodal models smaller, faster, and more useful — which usually means fighting with tokenizers, adapters, and whatever the latest efficiency trick is.

My current focus is on speech-vision alignment (getting a VLM to actually understand audio as more than just transcribed text) and efficient VLM fine-tuning (you don't need to update 7 billion parameters to teach a model new tricks).

Before all this, I spent a lot of time doing audio signal processing. Old habits die hard — I still think in spectrograms.


What I've been building

speech-vlm-fusion Cross-modal adapter toolkit: Whisper + LLaVA alignment
LightAdap Dynamic visual token pruning + LoRA, 2-3× speedup
audio-cap-bench Evaluation toolkit for audio captioning, with audio-specific metrics

Stack

Python PyTorch HuggingFace CUDA Linux LaTeX


Currently

  • Finishing experiments on audio-visual cross-modal alignment
  • Reading about speculative decoding (it's actually clever)
  • Trying to make audio-cap-bench useful for people who aren't me

Popular repositories Loading

  1. SkyCWO SkyCWO Public

  2. speech-vlm-fusion speech-vlm-fusion Public

    Cross-modal adapter toolkit for aligning speech encoders with vision-language models

    Python

  3. LightAdap LightAdap Public

    Dynamic visual token pruning for parameter-efficient VLM adaptation

    Python

  4. audio-cap-bench audio-cap-bench Public

    Benchmark and metrics toolkit for audio captioning and audio-language alignment evaluation

    Python

  5. PacketScope PacketScope Public

    Forked from Internet-Architecture-and-Security/PacketScope

    🎯 A general-purpose protocol stack analysis and debugging tool based on eBPF 🧰

    C

  6. Hive Hive Public

    Forked from JusperLee/Hive

    Python