Human narratives are formed upon rich sensory experiences encompassing not only vision and hearing but also temperature, wind, and bodily movement. However, existing large language model text generation research has primarily relied on static images o...
Human narratives are formed upon rich sensory experiences encompassing not only vision and hearing but also temperature, wind, and bodily movement. However, existing large language model text generation research has primarily relied on static images or text prompts, failing to adequately reflect the continuous sensor states and embodied experiences that real robots perceive. To address this limitation, this thesis proposes a sensor-conditioned embodied narrative generation framework that enables robots to automat- ically generate first-person literary narratives based on their own sensor states. To this end, we design a sensor encoder that projects a 12-dimensional sensor vector—comprising temperature, humidity, wind direction, 6-axis IMU, and relative angles—into high-dimensional embeddings, and fuse these into the input embedding space of a TinyLLaVA-based multimodal language model. Furthermore, we construct a synthetic sensor-text dataset of 40,000 pairs by combining virtual environments, weather conditions, and robot states, and fine-tune the model to generate first-person literary narratives corresponding to each sensor state. The proposed model receives sensor sequences collected from a real robot as input to gener- ate narratives, and we quantitatively evaluate performance across multiple dimensions—including overall quality, richness of embodied expression, and utilization of sensor states—using three independent LLM evaluators. Experimental results demonstrate that the proposed sensor-conditioned embodied narrative model consistently achieves superior preference ratings and scores compared to the pre-trained baseline model across most evaluation criteria. This research presents a novel form of embodied intelligence in which robots narratively “speak” their own sensory experiences, suggesting diverse future applications such as robot diaries, field experience documentation, and storytelling for human-robot interaction.