An Analytical Study of Diffusion-Based Approaches for Text-Guided Image Synthesis
Main Article Content
Abstract
Text-to-image generation using diffusion models has rapidly become a central focus in Generative Artificial Intelligence (AI), demonstrating significant improvements in image realism, semantic consistency, and controllability over earlier approaches such as Generative Adversarial Networks (GANs). These models are grounded in a robust probabilistic denoising framework, initially introduced through Denoising Diffusion Probabilistic Models (DDPM), which iteratively transform random noise into structured and meaningful images. Subsequent developments, including score-based generative modelling and extensive empirical validation, have further established diffusion models as the state-of-the-art in image synthesis. The incorporation of advanced language–vision encoders in systems like GLIDE, DALL·E 2, Imagen, and Stable Diffusion has enabled effective text conditioning, allowing precise alignment between textual prompts and generated visual content. Moreover, the introduction of Latent Diffusion Models has significantly reduced computational complexity by performing operations in compressed latent spaces, thereby enabling high-quality image generation on resource-constrained hardware. Recent innovations such as ControlNet, SDXL, and classifier-free guidance have further improved structural control, flexibility, and inference efficiency.
Despite these advancements and their growing adoption across creative, industrial, and accessibility domains, critical challenges persist, particularly concerning bias, authenticity, and ethical deployment. This paper presents a comprehensive analysis of the evolution, underlying architectures, and optimization strategies of diffusion-based text-to-image systems, and discusses future directions for developing efficient, interpretable, and responsible generative models
Article Details
Section
COPYRIGHT
Submission of a manuscript implies: that the work described has not been published before, that it is not under consideration for publication elsewhere; that if and when the manuscript is accepted for publication, the authors agree to automatic transfer of the copyright to the publisher.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work
- The journal allows the author(s) to retain publishing rights without restrictions.
- The journal allows the author(s) to hold the copyright without restrictions.
References
[1] T. Yin, P. Molchanov, and J. Kautz, "A new method for generating images with artificial intelligence," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Seattle, WA, USA, 2024.
[2] OpenAI, "Improvement of image generation through better captions: DALL-E system card 3," OpenAI, San Francisco, CA, USA, Tech. Rep., 2024.
[3] R. Podell et al., "An enhanced version of the latent diffusion model for high-resolution image generation," arXiv preprint arXiv:2307.01952, 2023.
[4] A. Zhang and J. Agrawala, "ControlNet: Adding conditions to conditional text/image diffusion models," arXiv preprint arXiv:2302.05543, 2023.
[5] T. Karras, M. Aittala, S. Laine, E. Harkonen, J. Hellsten, J. Lehtinen, and T. Aila, "Design of artificial intelligence-based generative models," in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), New Orleans, LA, USA, 2022.
[6] T. Salimans and J. Ho, "A faster sample collection method from a diffusion model," arXiv preprint arXiv:2202.00512, 2022.
[7] J. Ho and T. Salimans, "No classifier required when generating with diffusion models," arXiv preprint arXiv:2207.12598, 2022.
[8] C. Saharia et al., "Creating photo-realistic images from text with deep language processing techniques using a diffusion model," arXiv preprint arXiv:2205.11487, 2022.
[9] A. Ramesh et al., "Generating images through a hierarchical text conditioned latent variable model using CLIP," arXiv preprint arXiv:2204.06125, 2022.
[10] S. Gu, M. Chen, A. Radford, and K. Goyal, "Using a vector quantized diffusion model to create images from text," arXiv preprint arXiv:2111.14822, 2022.
[11] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, "High-resolution image synthesis with latent diffusion models," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), New Orleans, LA, USA, 2022.
[12] Google Research, "Imagen: Text-to-image diffusion model architecture and evaluation," Google Brain, Mountain View, CA, USA, Tech. Rep., 2022.
[13] Stability AI, "Stable diffusion model card and technical report," Stability AI, London, UK, Tech. Rep., 2022.
[14] A. Nichol et al., "GLIDE: Towards photorealistic image generation and editing with text-guided diffusion models," arXiv preprint arXiv:2112.10741, 2021.
[15] P. Dhariwal and A. Nichol, "Diffusion models beat GANs on image synthesis," in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), virtual, 2021.
[16] Y. Song, J. Sohl-Dickstein, D. Kingma, A. Kumar, S. Ermon, and B. Poole, "Score-based generative modeling through stochastic differential equations," in Proc. Int. Conf. Learn. Represent. (ICLR), virtual, 2021.
[17] J. Ho, A. Jain, and P. Abbeel, "Denoising diffusion probabilistic models," in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), virtual, 2020.