AI-based Personalization of Social Media Thumbnails Using the Stacked ID Embedding Method
DOI:
https://doi.org/10.34148/teknika.v14i2.1215Keywords:
Thumbnail, Stacked ID Embedding, Artificial Intelligence, Content Creator, Text-to-image GenerationAbstract
Social media content creators find it hard to make thumbnails shine on Instagram, YouTube, and TikTok, where graphic design skills happen to be a significant bottleneck. To address this issue, scientists developed a text-based image generation model, PotionPix, that allows users to generate thumbnails on the fly based on text prompts and relevant images via a "Stacked ID Embedding" method. This method combines multiple identity embeddings—e.g., user interests, platform context, and content genre—into one vector representation to guide the AI to create more personalized and contextually appealing thumbnails. The system integrates a diffusion-based image generator with the stacked embedding vectors to enable dynamic adaptation to different user intents. In tests, it was observed that how relevant and good the generated thumbnails were very much a function of how specific the input image was and how clear the prompt was. However, since the AI model used was not fine-tuned on the task of thumbnail generation specifically, the visual outputs sometimes were generic and lacked the strong call-to-action elements usually found in high-performing thumbnails. Despite this constraint, the usability test conducted with 120 respondents showed promising results—83.8% of the participants confirmed that PotionPix was indeed assistive in the thumbnail design process, particularly in terms of time and effort savings. The findings show the promise of AI-driven tools in enabling the democratization of design tasks for social media content creators, as well as suggesting future work in model fine-tuning for more domain-specific outcome.
Downloads
References
[1] C. I. Lestari and I. Irwansyah, ‘Kolaborasi Produksi Konten YouTube melalui Multi-Channel Network: Studi pada Kreator Sandy SS dengan Collab Asia’, J. Ris. Komun., vol. 4, no. 1, pp. 143–159, Mar. 2021, doi: 10.38194/jurkom.v4i1.152.
[2] M. Arar et al., ‘Domain-Agnostic Tuning-Encoder for Fast Personalization of Text-To-Image Models’, Jul. 13, 2023, arXiv: arXiv:2307.06925. doi: 10.48550/arXiv.2307.06925.
[3] B. Rahadjo, Belajar Otodidak Membuat Database menggunakan MySQL, 1st ed. Bandung: Bandung, Informtika, 2011. [Online]. Available: https://perpustakaan.binadarma.ac.id/opac/detail-opac?id=15408
[4] F. Firdausi, ‘Pengembangan Aplikasi Online Public Access Catalog (Opac) Perpustakaan Berbasis Mobile Pada Stai Auliaurrasyidin’, J. Intra Tech, vol. 4, no. 2, pp. 11–24, 2020, doi: https://doi.org/10.37030/jit.v4i2.74.
[5] C. Zhang, C. Zhang, M. Zhang, I. S. Kweon, and J. Kim, ‘Text-to-image Diffusion Models in Generative AI: A Survey’, Nov. 08, 2024, arXiv: arXiv:2303.07909. doi: 10.48550/arXiv.2303.07909.
[6] J. Liu and Y. Zhu, ‘Precise Correspondence Enhanced GAN for Person Image Generation’, Neural Process. Lett., vol. 54, pp. 5125–5142, 2022.
[7] Z. Li, M. Cao, X. Wang, Z. Qi, M.-M. Cheng, and Y. Shan, ‘PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding’, Dec. 07, 2023, arXiv: arXiv:2312.04461. doi: 10.48550/arXiv.2312.04461.
[8] J. D. Stemple, ‘Job Satisfaction Of High School Principals In Virginia’, Diss. Submitt. Fac. Va. Polytech. Inst. State Univ., Apr. 2004.
[9] R. Gal et al., ‘An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion’, Aug. 02, 2022, arXiv: arXiv:2208.01618. doi: 10.48550/arXiv.2208.01618.
[10] I. M. De La Vega Hernández, A. Serrano Urdaneta, and O. Schiappa-Pietra, ‘Scientific mapping of artificial intelligence as an emerging field of knowledge’, in Handbook of Research on Artificial Intelligence, Innovation and Entrepreneurship, E. Carayannis and E. Grigoroudis, Eds., Edward Elgar Publishing, 2023, pp. 8–28. doi: 10.4337/9781839106750.00008.
[11] J. Ho, A. Jain, and P. Abbeel, ‘Denoising Diffusion Probabilistic Models’, 34th Conf. Neural Inf. Process. Syst. NeurIPS 2020 Vanc. Can., Dec. 2020.
[12] M. Irfan, H. Siregar, and J. T. Handoko, ‘Pengembangan Dan Integrasi Aplikasi Prediksi Jumlah Gagal Produksi PC Menggunakan Metode Triple Exponential Smoothing Pada Sistem Aplikasi Produksi Di PT Tera Data Indonusa, Tbk’, Pros. Semin. Nas. Darmajaya, vol. 1, 2023.
[13] X. Jia et al., ‘Taming Encoder for Zero Fine-tuning Image Customization with Text-to-Image Diffusion Models’, Apr. 05, 2023, arXiv: arXiv:2304.02642. doi: 10.48550/arXiv.2304.02642.
[14] W. Peebles and S. Xie, ‘Scalable Diffusion Models with Transformers’, Mar. 02, 2023, arXiv: arXiv:2212.09748. doi: 10.48550/arXiv.2212.09748.
[15] A. Radford et al., ‘Learning Transferable Visual Models From Natural Language Supervision’, Feb. 26, 2021, arXiv: arXiv:2103.00020. doi: 10.48550/arXiv.2103.00020.
[16] J. Sutrisno and V. Karnadi, ‘Aplikasi Pendukung Pembelajaran Bahasa Inggris Menggunakan Media Lagu Berbasis Android’, J. Comasie, vol. 4, no. 6, 2021.
[17] I. S. Akbar and T. Haryanti, ‘Pengembangan Entity Relationship Diagram Database Toko Online Ira Surabaya’, Comput. Insight J. Comput. Sci., vol. 3, no. 2, pp. 28–35, Jul. 2023, doi: 10.30651/comp_insight.v3i2.12002.
[18] T. H. Trinh, M.-T. Luong, and Q. V. Le, ‘Selfie: Self-supervised Pretraining for Image Embedding’, Jul. 27, 2019, arXiv: arXiv:1906.02940. doi: 10.48550/arXiv.1906.02940.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Teknika

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.















