← All work - Problem
- Generating a photorealistic image from a sentence is a resolution problem as much as a semantic one. A single generator asked to go straight from text to a sharp image tends to produce something that is either blurry or ignores half the description.
- My role
- Sole implementer, as part of ArIES at IIT Roorkee.
- Approach
- Stacked the generation in stages rather than attempting it in one — an early stage fixes shape and colour from the text embedding, a later stage refines detail conditioned on both the text and the previous output. More moving parts to train and tune, but each stage has a tractable job.
- Outcome
- Trained on CUB-200-2011 — 200 bird species, 11,788 images — producing recognisable, text-faithful samples.
Why staged generation
The failure mode of single-stage text-to-image at this scale, and what
stacking buys you.
Conditioning
How the text embedding enters each stage, and what changes between them.
Results and honest limitations
Sample outputs, and where it falls down. The limitations section is what makes
this read as engineering rather than a course submission.