← Back to feed

Creo: From One-Shot Image Generation to Progressive, Co-Creative Ideation

L5 · ResearcherResearcharXiv· 4/15/2026

Directly addresses a fundamental HCI and generative AI challenge: how to design image generation systems that preserve user agency and creative intentionality rather than anchoring decisions prematurely.

AI Summary

Creo is a multi-stage text-to-image system that scaffolds image generation from rough sketches to high-resolution outputs, allowing users to maintain fine-grained control and creative agency throughout the process. The paper presents a comparative study showing that progressive generation with intermediate abstractions and decision-locking mechanisms increases user ownership, reduces output homogeneity, and improves controllability compared to one-shot generation.

Excerpt

Text-to-image (T2I) systems enable rapid generation of high-fidelity imagery but are misaligned with how visual ideas develop. T2I systems generate outputs that make implicit visual decisions on behalf of the user, often introduce fine-grained details that can anchor users prematurely and limit their ability to keep options open early on, and cause unintended changes during editing that are difficult to correct and reduce users' sense of control. To address these concerns, we present Creo, a mul

Read Original
0 upvotes · 0 downvotes · 1 min read

Related Articles

L5 · ResearcherResearcharXiv
When Personality Meets Quantization: A Layer-wise MBTI Analysis of Quantized LLMs

Researchers conducted a layer-wise MBTI personality analysis of quantized LLMs, finding personality is emergent and quantization-sensitive rather than static.

L5 · ResearcherResearcharXiv
Agentic Autoresearch for Cell-Edge Power Control: Radically Redefining the Researcher's Role

A research paper presents an autonomous AI agent that designed a machine learning algorithm for cell-edge power control in wireless networks, achieving strong performance with minimal human input.

L5 · ResearcherResearcharXiv
Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings

Google researchers present an autonomous AI system that synthesizes multimodal geospatial data and outperforms expert baselines across health, disaster risk, and food security prediction tasks.

L5 · ResearcherResearcharXiv
PlanSightRAG: A Visual-First Multimodal RAG for Automating Question Answering and Compliance Checking for Civil Standard Plans

Researchers developed PlanSightRAG, a visual-first multimodal RAG system that achieves 91.47% recall on zero-shot retrieval for civil engineering plan compliance checking.

L5 · ResearcherResearcharXiv
Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders

Researchers apply sparse autoencoders to interpret latent representations in a neutrino foundation model, discovering underutilized physical concepts that improve angular resolution from 20.2° to 3.2°.