From Prompting to Composing: A Spatial Canvas Interface for Poster Generation
Compo generates posters from a spatial canvas of bindings, improving compositional control over prompt-only systems.
The paper replaces text prompts for poster generation with a Spatial Canvas Interface using semantic, identity, text, and pixel bindings plus per-element and global text specifications. Compo is adapted from a pretrained image editing model and supports both direct canvas input and an agentic mode that turns a high-level request into a planned canvas. A synthetic supervision pipeline covers binding types and combinations, and experiments report stronger compositional control than general image models and dedicated poster systems.
- Users specify semantic, identity, text, and pixel bindings directly in 2D space.
- Compo is adapted from a pretrained image editor rather than trained from scratch.
- It supports explicit canvas construction and an agentic planning mode.
- A new benchmark measures single bindings and joint compositional adherence.
Full article198 words · extracted from huggingface.co · click to collapse
Text prompting is an indirect interface for poster generation, requiring users to encode inherently two-dimensional composition intent into a one-dimensional sequence of words. We introduce a Spatial Canvas Interface that enables users to directly compose generation intent in space through four complementary binding types: semantic, identity, text, and pixel, together with Text Specifications for individual elements and global appearance. Based on this interface, we develop Compo, a poster generation model adapted from a pretrained image editing model to understand Spatial Canvas inputs and Text Specifications. Compo supports both direct inference, where users explicitly construct the canvas, and agentic mode, where a high-level request is automatically translated into a planned Spatial Canvas. To train Compo, we develop a scalable pipeline that automatically constructs supervision data for different binding types and their combinations, enabling efficient adaptation without training a specialized poster generator from scratch. We further introduce a benchmark that evaluates adherence to individual binding types and their joint composition. Experiments show that Compo achieves stronger compositional controllability than both general-purpose image generation models and dedicated poster generation systems while maintaining high visual quality. By decoupling intent specification from visual generation, our work shifts poster generation from prompting toward composing.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.12230