Working on safety alignment of Text-to-Image & T2V diffusion and flow matching models using a neurosymbolic approach called scene graphs.
Developing a novel methodology to reduce hate content in T2I Diffusion models using a safety potential-guided rectified flow matching in the CLIP embedding space.
Benchmarked debiasing and safety methodologies — TRCE, CURE, SAEUron, and DoCo — on lab's DETONATE dataset.