Prova de Doutoramento do aluno Daniel Alexandre Pratas de Oliveira

Área: Engenharia Informática e de Computadores

Despacho de nomeação de Júri

Edital

Título da Tese: From Pixels to Prose: Visually Grounded Storytelling through Entity Linking

Local da Prova:  https://tecnico-pt.zoom.us/j/91720178234 

Data: 10/11/2026

Hora: 14h00

Abstract: This dissertation develops a model integrating language models with computer vision techniques to automatically generate stories from visual data, addressing four key challenges: explicit visual grounding, temporal consistency of detected entities, semantic alignment, and narrative evaluation. The work progresses from grounding individual visual elements to generating coherent multi-image narratives. We introduce GroundCap, a dataset of 52,016 images with captions maintaining object identity through explicit references. This establishes transparent mechanisms for aligning textual elements with visual representations via an ID-based system, enabling consistent object tracking and complex reference resolution. Alongside this dataset, we present PixtralGroundCap, a caption generation model trained on GroundCap. For multi-image contexts with temporal coherence, we developed StoryReasoning with 4,178 stories from movie scene sequences. This dataset incorporates structured scene analyses explicitly modeling characters, objects, settings, and narrative progression through tabular representations. The system maintains consistent entity IDs across images while explicitly linking narrative to visual elements, reducing hallucinations by 12.3% compared to non-finetuned models. Alongside this dataset, we introduce Qwen Storyteller, a story generation model trained on StoryReasoning. To improve entity re-identification across images, we introduced contrastive reinforcement learning using synthetic negative examples. This teaches models when to establish entity connections between images, improving re-identification and increasing the share of entities recognized in 5+ images by 13.7%. Alongside this, we introduce QwenStoryteller2, a story generation model trained with contrastive reinforcement learning. To address semantic hallucinations where models misinterpret character relationships and dialogue attribution, we created StoryMovie, a dataset of 1,757 stories integrating movie scripts and subtitles for ground-truth semantic context. The script-subtitle alignment provides authentic character names, relationships, and dialogue attribution. Alongside this dataset, we introduce QwenStoryteller3, which outperforms previous models on subtitle alignment.

Tópicos: