An Analytical Study of Diffusion-Based Approaches for Text-Guided Image Synthesis

Main Article Content

Himangi Ahuja

Abstract


Text-to-image generation using diffusion models has rapidly become a central focus in Generative Artificial Intelligence (AI), demonstrating significant improvements in image realism, semantic consistency, and controllability over earlier approaches such as Generative Adversarial Networks (GANs). These models are grounded in a robust probabilistic denoising framework, initially introduced through Denoising Diffusion Probabilistic Models (DDPM), which iteratively transform random noise into structured and meaningful images. Subsequent developments, including score-based generative modelling and extensive empirical validation, have further established diffusion models as the state-of-the-art in image synthesis. The incorporation of advanced language–vision encoders in systems like GLIDE, DALL·E 2, Imagen, and Stable Diffusion has enabled effective text conditioning, allowing precise alignment between textual prompts and generated visual content. Moreover, the introduction of Latent Diffusion Models has significantly reduced computational complexity by performing operations in compressed latent spaces, thereby enabling high-quality image generation on resource-constrained hardware. Recent innovations such as ControlNet, SDXL, and classifier-free guidance have further improved structural control, flexibility, and inference efficiency.


Despite these advancements and their growing adoption across creative, industrial, and accessibility domains, critical challenges persist, particularly concerning bias, authenticity, and ethical deployment. This paper presents a comprehensive analysis of the evolution, underlying architectures, and optimization strategies of diffusion-based text-to-image systems, and discusses future directions for developing efficient, interpretable, and responsible generative models


Article Details

Section

Articles

Author Biography

Himangi Ahuja

Department of Computer Science and Engineering

Geetanjali Institute of Technical Studies

References

[1] T. Yin, P. Molchanov, and J. Kautz, "A new method for generating images with artificial intelligence," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Seattle, WA, USA, 2024.

[2] OpenAI, "Improvement of image generation through better captions: DALL-E system card 3," OpenAI, San Francisco, CA, USA, Tech. Rep., 2024.

[3] R. Podell et al., "An enhanced version of the latent diffusion model for high-resolution image generation," arXiv preprint arXiv:2307.01952, 2023.

[4] A. Zhang and J. Agrawala, "ControlNet: Adding conditions to conditional text/image diffusion models," arXiv preprint arXiv:2302.05543, 2023.

[5] T. Karras, M. Aittala, S. Laine, E. Harkonen, J. Hellsten, J. Lehtinen, and T. Aila, "Design of artificial intelligence-based generative models," in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), New Orleans, LA, USA, 2022.

[6] T. Salimans and J. Ho, "A faster sample collection method from a diffusion model," arXiv preprint arXiv:2202.00512, 2022.

[7] J. Ho and T. Salimans, "No classifier required when generating with diffusion models," arXiv preprint arXiv:2207.12598, 2022.

[8] C. Saharia et al., "Creating photo-realistic images from text with deep language processing techniques using a diffusion model," arXiv preprint arXiv:2205.11487, 2022.

[9] A. Ramesh et al., "Generating images through a hierarchical text conditioned latent variable model using CLIP," arXiv preprint arXiv:2204.06125, 2022.

[10] S. Gu, M. Chen, A. Radford, and K. Goyal, "Using a vector quantized diffusion model to create images from text," arXiv preprint arXiv:2111.14822, 2022.

[11] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, "High-resolution image synthesis with latent diffusion models," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), New Orleans, LA, USA, 2022.

[12] Google Research, "Imagen: Text-to-image diffusion model architecture and evaluation," Google Brain, Mountain View, CA, USA, Tech. Rep., 2022.

[13] Stability AI, "Stable diffusion model card and technical report," Stability AI, London, UK, Tech. Rep., 2022.

[14] A. Nichol et al., "GLIDE: Towards photorealistic image generation and editing with text-guided diffusion models," arXiv preprint arXiv:2112.10741, 2021.

[15] P. Dhariwal and A. Nichol, "Diffusion models beat GANs on image synthesis," in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), virtual, 2021.

[16] Y. Song, J. Sohl-Dickstein, D. Kingma, A. Kumar, S. Ermon, and B. Poole, "Score-based generative modeling through stochastic differential equations," in Proc. Int. Conf. Learn. Represent. (ICLR), virtual, 2021.

[17] J. Ho, A. Jain, and P. Abbeel, "Denoising diffusion probabilistic models," in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), virtual, 2020.