A Photo to Generative AI Workflow for Rapid 3D Heritage Representations
9th International Conference on Intelligent Computing and Communication, ICICC 2026, Hybrid, Delhi, Hindistan, 6 - 07 Şubat 2026, cilt.2021 LNNS, ss.296-305, (Tam Metin Bildiri)
- Yayın Türü: Bildiri / Tam Metin Bildiri
- Cilt numarası: 2021 LNNS
- Doi Numarası: 10.1007/978-3-032-28310-8_21
- Basıldığı Şehir: Hybrid, Delhi
- Basıldığı Ülke: Hindistan
- Sayfa Sayıları: ss.296-305
- Anahtar Kelimeler: AI-based modeling, Architectural digital twin, Cultural heritage documentation, Diffusion model, Gemini Flash 2.5, Hunyuan3D 2.1, Image-to-3D generation, LOD3 digital building representations, PBR textured mesh, Qwen Image Edit, Single-image 3D reconstruction
- İstanbul Üniversitesi-Cerrahpaşa Adresli: Evet
Özet
This paper investigates generative AI workflows that reconstruct 3D architectural heritage models from a single photograph. We compare two pipelines that combine a 2D image model with Tencent’s Hunyuan3D 2.1 image-to-3D system: (1) Gemini Flash 2.5 + Hunyuan3D and (2) Qwen-Image-Edit + Hunyuan3D. Using a dataset of 16 Turkish architectural landmarks photographed from sub-optimal viewpoints, each method first generates an isometric or gently re-angled view and then produces a PBR-textured GLB mesh approximating LOD3 building detail. The pipelines are evaluated in terms of visual fidelity of intermediate images, geometric completeness and sharpness of the 3D meshes, dimensional consistency, processing time, user experience, and downstream compatibility with BIM, GIS, VR and web viewers. Results show that both workflows deliver photorealistic, lightweight 3D assets within minutes, dramatically lowering the cost and expertise barrier compared with conventional photogrammetry or laser scanning. The Qwen-based pipeline better preserves original textures and colors, enriches side-facade information, and offers higher reproducibility thanks to its open-source model and public interfaces. The Gemini-based pipeline provides cleaner, stylized views but is constrained by closed access. We conclude by positioning single-image generative AI as a rapid, complementary tool for cultural heritage visualization, education and preliminary digital twin creation rather than a substitute for metric survey.