Special issue: Investigating the Relationship between Generative Artificial Intelligence and History Education
Introduction
In 2024, Google’s AI image generator Gemini caused controversy when it produced quite absurd historical representations, such as Black medieval Vikings and monks, Indigenous US founding fathers, female popes and Asian Wehrmacht soldiers. These blunders were apparently caused by an automatic modification of the prompts used, which was intended to make the desired images more ethnically and gender diverse, and thus counteract distortions in the predominantly White, male and European training corpus (Förtsch, 2024). The embarrassing incidents sparked discussions about political incorrectness, historical misrepresentation and the inherent tension between manipulation and historical integrity. In view of the rapidly growing capabilities of AI image and video generators to produce convincing historical ‘sources’, questions relevant to history education arise: what logic do these models follow? What criteria do they use to select their content? What distortions, biases and filters are at work? What perspectives (for example, Western or Eurocentric) are adopted according to the training corpora? Can their strengths and weaknesses also be used in history lessons to encourage historical thinking? For example, could they be used to help students develop a better understanding of historical concepts?
There is currently a lack of empirical research dedicated to answering such questions. The relevant research context for historical AI visualisations is characterised by interdisciplinary work on bias amplification, although empirical studies remain rare. Studies regularly confirm systematic stereotype reinforcement in large language models (LLMs) (for example, UNESCO, 2024) and image generators (for example, Bianchi et al., 2022; Furman et al., 2024; Hellmann, 2025), showing that text-to-image models disproportionately amplify demographic and cultural biases. However, these distortions appear to be gradually diminishing through expansions in the material base, at least in the case of large models (Vice et al., 2025). Whether the models demonstrate creativity in the sense of generating original ideas is debatable (Atkinson and Barker, 2023); rather, they generate hallucinatory combinations of existing image data, leading to style fusions and blends of historical and popular cultural elements in historical representations. Like LLMs, they operate probabilistically and therefore reinforce existing popular conceptions of history (Simanowski, 2025), while also reproducing the biases of the past (Bendel, 2025; Bianchi et al., 2022; Nicoletti and Bass, 2023; Thomson and Ryan, 2023). This naturally gives rise to problems, especially since young people are quite open-minded and trusting of AI-generated images in educational contexts (Pyae, 2024). However, these problems can also be made explicit so that pupils improve their historical understanding by being confronted with heavily biased or erroneous AI images (Madshaven et al., 2025; Szelenyi, n.d.).
Visual media already dominate information acquisition among students (MPFS, 2025). Soon, AI-generated images and films will also become a natural part of their media experience and their teaching materials, especially since they incur hardly any costs to produce. Before AI image generators rapidly find their way into school contexts, their modes of operation and problems, as well as their didactic possibilities, should be explored from a history education perspective. This is where the following study comes in. The article first provides a brief overview of the relevant theoretical and methodological references, then outlines the study’s design and results and finally suggests ways in which AI images could be used in the classroom.
Conceptual references
Historical theory distinguishes between historical sources and historical representations. The former are remnants from the past that we can use for reconstruction processes, while the latter are narratives about the past. Although AI-generated images are probabilistically composed from visual historical sources, they are treated in this study as historical representations, insofar as this composition is arranged according to an underlying narrative logic that responds to historically motivated user input. Whether AI demonstrates genuine historical judgement in generating historical images constitutes the central research question. To address this, the study employs three historical concepts whose visual outcomes can be effectively analysed through surface features:
-
historical significance: this is understood as the capacity to weigh persons and events perspectivally. This concept is regarded in Anglophone scholarship as central to historical thinking. Although reflexive criteria for determining significance are extensively debated in that tradition (for example, Cercadillo, 2001; Hunt, 2003; Counsell, 2004; Lévesque, 2005; Partington, 1980; Phillips, 2002; Seixas and Morton, 2013), these are not addressed in the analysis as they cannot be recognised in the surface features of the images. Instead, the analysis focuses on whether the models address specific events and individuals and, if so, whether these are historically and culturally relevant.
-
epochal authenticity: this concept targets the temporal plausibility of representations and can be analysed through deconstruction frameworks, such as those proposed by Jörn Rüsen (1994, 2013) and Hans-Jürgen Pandel (1993, 2006) (see also Gundermann, 2018). The study follows Pandel’s notion of representational and typological authenticity, as this seems more suitable for AI-generated image combinations without proper historiographical claims – it allows for more degrees of freedom within the framework of the ‘historical potential’ and the typicality of representations of the past than Rüsen’s approach of empirical triftigkeit (validity), which is more strictly focused on primary source proximity.
-
synchronic comparison: an evaluative juxtaposition that permits assessment across variants. Hartmut Kaelble’s (1999) work provides a suitable theoretical and operational basis for this aspect.
Additionally, the study adopts a historical-psychological perspective on training data. Just as individual historical consciousness rests on multiple layers of unconscious content, generative models recombine patterns embedded in training data that reflect what can be understood as the collective unconscious in the sense of Jung (1980) – the accumulated visual and narrative structures that also shape our historical culture (Ammerer, 2022; Chiru, 2024). The study examines archetypal hero representations (Campbell, 2008; Franz, 1994; Jung, 1980; Neumann, 1984) to identify which symbolic patterns from collective memory recur in generated historical imagery. This analysis reveals how training data – saturated with these archetypal structures – shapes what configurations of the past appear visually plausible to the model.
In terms of media analysis, the model of standardised visual content analysis developed by Geise and Roessler (2012) is adopted. This model is designed for the systematic and intersubjectively comprehensible examination of large bodies of communicative material for similarities and differences (Früh, 2007; Rose, 2016; Rössler, 2010). Following the scheme of image analysis developed by Erwin Panofsky (1975, 1979, 1997), which has also been adopted in the field of history education (Ammerer, 2019; Pandel, 2015), three levels of analysis can be identified. The first dimension comprises the objectively tangible surface features, such as formal design, perspectives, colours, concrete objects and people, and concrete situations, movements and relations. The second dimension comprises the latent elements of the internal structure, such as symbols and traditional signs, typical and stereotypical image content, and value implications. The third dimension comprises the deep structure of the image, or inherent constructions of meaning. The latter dimension contains the ‘actual meaning’ of the image (Panofsky, 1975), revealing the underlying values, basic attitudes and convictions that were typical of the time, as well as the context of reception and impact. Although this dimension is typically central to iconological image analysis, it will not be the focus of this study, as AI-generated images lack such intentional layers of (historical) meaning. Like current LLMs, image generators recycle (visual) patterns and create narrative probability formations based on existing interpretations (Burkhardt and Neubert, 2024).
Consequently, Panofsky’s framework is operationalised here within the logic of standardised content analysis: the codebook captures the first two dimensions (surface features and latent symbols) through formal and thematic categories, while the third dimension (iconology) serves not as a coding variable, but as a theoretical lens for interpreting the aggregated findings. Therefore, the focus lies on measurable structural elements processed using qualitative methods (for example, Mayring, 2010) established in history education research.
Analytical approach
This study examines the ability of AI image generators to solve tasks based on defined historical concepts, producing plausible visualisation results. It focuses on the key events and individuals that are prioritised in AI-generated historical images and whether this selection demonstrates an understanding of historical significance. Next, it considers the type of elements typical of a specific time that can be found in AI-generated historical images and whether this selection considers the concept of epochal authenticity. It also explores the ideas found in AI-generated images of contrasting historical lifeworlds and the historical logics of comparison, judgement and evaluation revealed in such images. Finally, it examines which historicising modes of representation can be found for a specific archetypal narrative motif in AI-generated images.
Between November and December 2024, the author generated 1,024 images as a dataset via web interfaces and desktop applications such as Automatic1111. Several basic models and derivatives were considered for creation. For the purposes of this study, the survey was limited to six common models which covered various training corpora:
-
Midjourney v.6.1 (trained on an art-centred dataset; 256 images).
-
Stable Diffusion (trained on larger, publicly accessible datasets such as LAION-5B; 64 images in Krea.ai and 128 images in Dreamshaper as a modified derivative).
-
Flux (trained on curated public datasets; 128 images in Creative Fabrica Flow and 32 images in FluxSchnell).
-
DALL-E 3 (trained on Microsoft/OpenAI proprietary datasets; 224 images).
-
Leonardo Phoenix (possibly trained on multimodal corpora including Canvas data; 128 images).
-
Firefly AI (trained on Adobe Stock images and royalty-free works; 64 images).
The sample composition does not reflect the market shares of the respective models at the time of the survey. Reliable estimates and comparative user statistics on AI image generation were unavailable. Midjourney was assumed to dominate the market due to its highly artistic quality, while DALL-E, Flux and Stable Diffusion were assumed to be popular due to their easy availability and flexibility. Other models were expected to occupy niches (Borji, 2022). Most models used in 2024 were based on the diffusion technique, whereby more and more noise is added to an image during training so that AI can recognise schematic structures. Then the noise is removed so that AI can learn to generate objects independently with the help of a text transformer (Chen et al., 2024; Luo, 2022).
Eight prompts in English were used, each generating 128 images. These were fine-tuned in iteration loops before the dataset was created, ensuring that the models would generate images that corresponded to their intended purpose. Generally, the instructions were kept as concise as possible to minimise influence on the content.
Almost all models implemented the prompts without any externally recognisable text modulation. Only Leonardo Phoenix used a visible multimodal approach in the interface to automatically improve prompts, using an LLM to make general prompts more specific. For instance, the prompt ‘Generate a picture showing a pivotal moment in history, concrete situation, artistic depiction’ was transformed into the following detailed instructions within the model framework:
Create a highly detailed and vibrant illustration capturing a pivotal moment in history, such as the signing of the Magna Carta or the first landing on the moon, showcasing a concrete historical event, inspired by classic historical artistic depictions like those of Delacroix or David, with a focus on dramatic lighting, intense emotions, and intricate textures, featuring prominent historical figures with distinct facial features, expressive skin tones, and period-specific attire, set against a richly colored, atmospheric background that evokes the sense of grandeur and importance of the moment, with bold brushstrokes, ornate details, and a sense of movement and energy, all carefully composed to draw the viewer’s eye to the central action of the scene.
The material was compiled according to the research questions and the individual images were treated as coding units. Formal characteristics were recorded and 99 analysis categories were defined during the initial review. Subsequently, 359 subcategories were defined through inductive iteration. Coding was carried out in Excel using a codebook, with 8,670 dichotomous, trichotomous, polytomous or multiple-choice codes assigned. The subcategories were operationalised across four dimensions (historical significance: 124; epochal authenticity: 170; synchronic comparison: 32; archetypal representation: 33) with defined decision rules. Decision rules for each code were defined operationally. For example, Code 1A_1_b (Event – Generic) was assigned when (a) no identifiable historical referent was discernible, (b) temporal specificity was diffuse or unlocalised and (c) the depicted scene lacked concrete landmarks or recognisable figures. The complete codebook is deposited in https://doi.org/10.17605/OSF.IO/5W8NH. The category system was tested for inter-rater reliability using a sample of almost 10% of the total dataset, yielding an average Kappa value of κ = 0.84, generally considered to indicate good agreement (Krippendorff, 2004; Kuckartz, 2016; Landis and Koch, 1977). Finally, the coded data was aggregated and interpreted in relation to the research questions to identify frequencies, patterns and structures.
Findings
Historical significance
The first question of interest was which key events and people are prioritised in AI-generated historical images and whether this selection adheres to a scholarly tradition or reflects concepts related to significance. To explore this aspect, images generated from the two general prompts targeting ‘pivotal moments’ and ‘famous persons’ were analysed first (see Table 1). As these initial outputs often lacked clear historical localisation, a third, more specific, series was created focusing on national history, using the prompt: ‘The single most important historical event in German history.’ In all three image series, it was striking that AI was virtually unable to refer to specific people or events: 65% of general events, 79% of German events and 82% of persons were purely generic and abstract, that is, they appeared historically plausible, but could not be assigned to any specific historical facts or personalities (see Figure 1 and Figure 2).
Experimental prompts
| Aspect | Prompt | No. |
|---|---|---|
| Historical significance | ‘Generate a picture showing a pivotal moment in history, concrete situation, artistic depiction.’ | 1A |
| ‘Generate a picture showing a famous person from history, specific historical person, artistic depiction.’ | 1B | |
| ‘The single most important historical event in German history.’ | 1C | |
| Epochal authenticity | ‘Generate a picture showing a typical scene from the historical period between 500 and 1500 ad.’ | 2A |
| ‘Generate a picture showing a typical scene from the historical period between 1500 and 1900 ad.’ | 2B | |
| ‘Generate a picture showing a typical scene from the historical period from the 20th century.’ | 2C | |
| Synchronic comparison | ‘Generate a split-screen image illustrating the different life in East and West Germany during the Cold War.’ | 2D |
| Archetypal representation | ‘Generate an allegorical (human) representation of individual heroism and bravery.’ | 3A |
A further fifth of global event depictions, 9% of German history depictions and 5% of people depictions exhibited a fantastical and diffuse character. References to time, people, events, reality and fantasy were mixed, particularly in DALL-E images. Plausible elements appropriate to the subject matter were regularly mixed with references and motifs (for example, religious) unrelated to the event. Events that could be traced back to a specific time and place were usually modelled on iconic references (see Figure 3 for an example) and could be identified via AI-supported image searches, such as Google Lens.
In some cases, such as the proclamation of the German Republic in 1918, several events are blended into one scene, including the German October uprising of 1923, the Russian October Revolution of 1917 and the German Spartacist uprising of 1919. An AI prompt in Leonardo Phoenix instructing the software to depict the signing of the Magna Carta resulted in a Baroque-style image that included elements of space travel (see Figure 4). When Claude 3.7 analysed this composition, it identified a different event: the Peace Treaty of Utrecht in 1713. In most cases, depictions of people simply diffused and recombined digitally available portraits of historical figures, which often made identification difficult.
Identifiable events included the first moon landing, the fall of the Berlin Wall, the Battle of Leipzig, the storming of the Bastille and Washington’s crossing of the Delaware. Leonardo Phoenix created most of the images for these events, using prompt enhancement to give the image generator more precise instructions in terms of content and aesthetics (see Figure 5). Midjourney also produced convincing results for the clearly identifiable historical figures (Albert Einstein, Marie Curie, Leonardo da Vinci, Abraham Lincoln, Napoleon Bonaparte and Cleopatra), as its training corpus apparently contains substantial portraits of important historical figures.
In the case of German history, national history is only recognisable in just over a fifth of the images; for the most part, the depiction is unspecific and in 9% of cases it shows clear anachronisms. Relevant landmarks include iconic German cathedrals (15 times), the Brandenburg Gate (11 times), the Berlin Wall (four times) and the Reichstag building (three times). At a symbolic-generic level, flags (68%), soldiers (56%) and churches (32%) mark the images as historical. A striking feature of the images is their violent nature: almost half depict scenes of war and marches, followed by public gatherings and scenes of riots (27%) and revolution (20%). The images also contain a high number (23%) of depictions of fires and conflagrations. In terms of atmosphere, around half of the images can be categorised as ‘dramatic’. Most images depicting German history come from the modern era (especially the late nineteenth and twentieth centuries), while scenes of general history are mostly pre-modern. Here, too, scenes of war and turmoil dominate, with a disturbing atmosphere, and in more than a quarter of cases, the depiction is so destructive that one could speak of downright apocalyptic scenes. In both series of images, only a few individuals are at the centre; usually, historical agency lies with more or less amorphous groups.
Notably, historical figures can almost exclusively be categorised as belonging to Western history (91%), predominantly European. Most people (38%) appeared to be from the late nineteenth or twentieth century, while a quarter could not be categorised chronologically. There was no clear focus on professional categories (for example, politicians, inventors or artists); rather, most of the depictions were unspecific. Unsurprisingly, the vast majority of figures were male (83%). In terms of style, all three series of images were predominantly executed as paintings or drawings, in keeping with the training material.
So, how can the historical judgement skills of image AI be evaluated? The results are hardly convincing. Identifiable events and personalities are sparse and can be considered canonical, with a historically significant nature that is open to debate (see, for example, Workman Publishing, 2016; Zimmermann, 2015). Rather than processing the request historically, the models seem to resort to appropriately titled or annotated training sources, such as historical paintings. Prompt-generating text-image AI provides the best results, but it can be assumed that it has used existing lists of examples, as current reasoning models such as o3-mini or Claude 3.7 Sonnet do when referring to relevant websites in their results.
To better understand AI’s construction logic, an advanced reasoning model (Claude 3.7 Sonnet) visually interpreted the images of the first series. AI was asked to judge the images based on whether the prompt was implemented comprehensibly and whether a historically significant event was clearly recognisable. Seven images were rejected due to explicit content. For the remaining images, the model matched the human coding completely or approximately in 93% of cases. However, in nine cases, AI identified specific events in images that the author had coded as generic, referencing embedded symbols. These included the Great Fire of London in 1666, inferred from a flame-distorted metropolis with Big Ben (erected much later); the destruction of Jerusalem in 70 ad inferred from burning ancient temples from various cultural areas; the fall of Rome in 410, inferred from ancient and medieval battle scenes; and the 1929 stock market crash, inferred from desperate crowds in front of historicist-looking US buildings. Although AI pointed out the lack of clear landmarks and anachronistic/atopic pictorial elements, it justified its choices by referencing the ‘essence’ of the events.
Epochal authenticity
Regarding epochal authenticity, three series were generated using prompts that called for ‘typical scenes’ from distinctive periods: the Middle Ages (500–1500 ad), the early modern/modern era (1500–1900 ad) and the twentieth century (see Table 1). It should be first noted that all the models created abstract scenes as intended, which were not related to identifiable events or individuals. In the series depicting the Middle Ages and the early modern and modern eras, the scenes were consistently realistic and homogeneous in terms of the era depicted (with hardly any anachronisms). In contrast, those depicting contemporary history also contained fantastic and grotesque narrative elements, primarily generated by Midjourney, which suggest increased influences from literature, including science fiction (see Figure 6).
The epochal appearance was convincing overall, but unclear at the temporal edges. In the medieval series, two-thirds of the paintings had a medieval or early modern appearance, while a quarter could only be dated diffusely to the pre-modern period (see Figure 7). In the modern series, 70% of the paintings also had a medieval or early modern appearance, while only 14% could be assigned to the period between the Baroque era and the late nineteenth century (see Figure 8). Among the twentieth-century pictures, 40% had a diffuse appearance between the fin de siècle and the First World War, 17% could be assigned to the interwar period and the Second World War, and a third to the second half of the century. Leonardo Phoenix consistently temporalised the images plausibly in all three series, while Firefly almost consistently produced images that could not be classified in terms of time.
Geographically speaking, 63% of the medieval scenes had a central or northern European flavour, while 28% had a diffuse Mediterranean feel. Almost 40% of the pictures showed unplastered stone buildings; almost a third showed castles; 25% showed churches, temples and mosques; and 12% showed ruins or Central European half-timbered houses. Of the people depicted, 94% wore long robes. In the early modern and modern images, 79% were clearly attributable to Europe and 15% to the Mediterranean world (European/Islamic), with American, Asian and African settings omitted. Almost all of the people depicted wore long dresses and skirts, and historicising building elements such as half-timbered houses (42%), unplastered stone buildings (31%) and churches and castles (27% each) stood out. A diffusely ‘Western’ setting also dominated contemporary history, with a tendency towards US settings (for example, Figure 9). Here too, the genuinely Asian, African and Islamic worlds were omitted. The typical historical elements of this period were mainly cars (46%), advertising signs and leisure activities (31%), as well as tall concrete buildings. People were frequently depicted wearing hats (47%) and photorealism was sometimes employed (15%). More often than in the other two periods, the scene was set in urban areas (83%).
There is a noticeable progression in gender representation over time: while nearly half of the medieval pictures showed men exclusively, this number dropped substantially in subsequent periods. Of twentieth-century pictures, three-quarters were gender-balanced.
In terms of scenes depicted, both medieval and modern paintings were dominated by calm depictions of everyday life and market scenes, followed by assembly scenes and war scenes. Contemporary images were dominated by depictions of busy shopping streets (28%), followed by unspecific everyday scenes (23%) and misanthropic concrete jungles, which were often accompanied by a stimulating (30%) or dreary (22%) atmospheric design.
Overall, the depiction of the eras can largely be regarded as plausible, even though the perspective is almost exclusively Western and rather generous with regard to the boundaries of the era. Urban life was mostly depicted alongside generic historical references in terms of buildings, clothing and interactions. It is likely that the depictions were based on genre painting and photographic evidence. Therefore, it can be assumed that no genuinely historical logic aimed at typicity was employed, except (in part) for the prompt-generating text-image AI.
Synchronic comparison
To test AI’s ability to compare synchronic living environments, the East–West conflict in Germany (1945–90) served as a case study. The corresponding prompt explicitly requested a ‘split-screen image illustrating the different life in East and West Germany during the Cold War’ to force a direct juxtaposition. It should be noted that the models generally succeeded in displaying the correct time period and everyday scenes (although these were sometimes highly condensed). However, several models (DALL-E, Flux and Dreamshaper) struggled with the split-screen presentation method, creating continuous images divided only by colour nuances (see Figure 10). Half of the pictures appeared to be drawn, half photorealistic.
The terminology of the ‘Cold War’ was reflected in wintry scenes (38%), with a dreary atmosphere (52%) interspersed with images of military personnel (19%), facilities (9%), faceless concrete buildings (38%) and war ruins (16%) being particularly noticeable (see Figure 11).
The symbolic reference to Germany was primarily demonstrated through flags (30%), landmarks such as the Brandenburg Gate (9%) and the Berlin Radio Tower (6%), and recognisable car brands. However, almost all models failed to make a real comparison: in 88% of pictures, no distinction was made between East and West Germany, nor was any evaluation of the two living environments recognisable. No historical judgement was made. Only the Leonardo Phoenix model with its AI-enhanced prompting was clearly judgemental, though in a rather unnuanced way:
A split screen image contrasting the stark differences between East and West Germany during the Cold War era, with the left side depicting a gloomy, rundown East German cityscape under communist rule, featuring a gray concrete apartment complex with crumbling walls, a dimly lit streetlamp, and a lone figure in a long coat walking away from the viewer, set against a bleak, overcast sky, while the right side showcases a vibrant, prosperous West German city scene, with a well-lit shopping street lined with colorful storefronts, bustling with people of diverse ages and styles, including a woman with curly blonde hair and a bright smile walking towards the viewer, set against a clear blue sky with a few wispy clouds, highlighting the vastly different living standards and atmospheres of the divided nation.
Accordingly, dystopian ‘rubble worlds’ were juxtaposed with utopian ‘leisure societies’, with economic, freedom, progress and community differences being figuratively expressed (see Figure 12 and Figure 13).
This aspect of the study highlights the fact that image AI does not yet display historical thinking skills. Historical judgements can only be found in the products of generative text AI – which probably does not reason independently as well, but rather implements internet search results that reproduce existing master narratives on a topic.
Archetypal modes of representation
Finally, to investigate whether archetypal motifs assume historicised forms in AI imagery, the model was instructed to generate an ‘allegorical (human) representation of individual heroism and bravery’. The results revealed an absence of contemporary notions of heroism, such as firefighters or astronauts, and instead showed a prevalence of mythical-supernatural (48%) and historicised (52%) representations, the latter of which were almost exclusively associated with ancient and medieval times. This is hardly surprising in a post-heroic society (Münkler, 2006), where heroism is both fictionalised and historicised. There were clear distortions in terms of ethnicity (95% of the people were White or European) and gender (80% male). Only two models (DALL-E and Stable Diffusion) created female or androgynous heroes (see Figure 14), suggesting the existence of pre-filtering.
Heroism was predominantly expressed in a dramatic manner. Martial poses (see Figure 15) and inspiring gestures (see Figure 16) were common, but stoic and tragic heroism were also depicted.
Symbolically, heroism was represented with visual attributes such as capes and fluttering capes (87%), weapons (56%), naked muscular torsos (42%), long manes of hair (39%), solar and celestial symbols (37%), wings (18%), rocks and conquered mountains (19%) and halos (12%). In many cases, these symbols merged (see Figure 17).
Each generative model focused on different aspects. DALL-E presented diverse cultural and ethnic portrayals of heroism, while Firefly depicted powerful fighters in a video game style. Leonardo Phoenix showcased superhuman war leaders, and in the Stable Diffusion and Flux models, angelic hermaphrodites posed alongside generals. Drawing on its art-historical expertise, Midjourney depicted pathos-laden male strength and tragic heroism, probably inspired by nineteenth-century European national romantic iconography, with its dramatic viewpoints, pathetic gestures and flag dynamics.
Conclusion and outlook
AI image generators compile historical visualisations from existing training data, which can appear quite plausible. However, they are not yet based on independent historical thinking and judgement skills. In the present study, this was primarily evident in the tasks on historical significance and evaluative historical comparison. For the most part, no identifiable events, personalities or comprehensible comparisons of lifeworlds could be recognised. Generic depictions of typical period scenes were more successful, although these were probably based less on an understanding of epochal authenticity and more on a successful combination of historical pictorial references (for example, symbols, landmarks, scenery and costumes).
The differences between the models were remarkable. Midjourney, which had evidently been strongly trained in paintings from European art history, appealed with its high artistic standards, but tended towards particularly dramatic, warlike and sometimes grotesque depictions. DALL-E also displayed a tendency towards fantastical exaggeration, but was much more contemporary in style and made a recognisable effort to achieve ethnic and cultural diversity. Stable Diffusion and Flux derivatives were simpler in style, whereby the rather generic compilation style made the depictions of significance and historical comparison less convincing. Firefly proved to be hardly applicable for historical representations due to the data basis used. The most plausible and concrete results overall for the three genuinely historical tasks were delivered by Leonardo Phoenix, a model that uses prompt chaining for an AI cascade process. With multimodal applications such as this, there appears to be great potential for further development towards a closed AI cycle in which the image result is checked and improved in an iterative feedback loop process. As has been stated many times before (for example, Bianchi et al., 2022; Breithut, 2022; Thomson and Ryan, 2023), the models amplified numerous data-based stereotypes, displaying a clear bias towards a Western/European, urban, middle-aged, male perspective. These provide good starting points for reflection on history in lessons.
Future research could include longitudinal studies on the development of historical representation accuracy, cross-cultural comparisons of different AI training corpora and experimental designs investigating the effects on recipients. As the development, integration and self-examination of the models progresses, research approaches need not be limited to objectifiable, manifest and semi-latent characteristics; they could also focus on the deep structure of images. For example, iconographic comparative studies could be conducted with historical models, and structural analyses could be performed on anachronisms, anachronistic hybrid forms and collective distortions.
Transfer into teaching practice
AI-generated images are set to become a common feature in history lessons. However, the findings emphasise the importance of critical reflection alongside this development. To use AI images to promote historical thinking, a suitable methodology is required and should be incorporated into teacher training.
Crucially, the didactic potential of AI-generated images does not depend on future, more elaborated models. Rather, it is precisely the current probabilistic nature of these models that provides useful learning opportunities. When students use AI to visualise higher-order thinking concepts, they can expose the model’s reliance on visual correlations (for example, fire = revolution, ruins = war) and critically discuss why such pattern-matching falls short of genuine historical reasoning. The limitations of the technology thus become the starting point for reflection.
This could include creating and analysing images that implement historical concepts such as causality, perspective or morality. The dataset used in the study could be used to test the application of AI to historical didactic criteria for significance, sensitise students to anachronisms or practise the quality criteria and methodological steps of historical comparisons. Authentic-looking AI images can be used to teach historical source criticism, for example by verifying their origin using an image search.
Conversely, students can use AI to creatively express events and people whose significance they have identified using traditional teaching methods (for example, Bradshaw, 2006; Historical Thinking Project, n.d.; Kitson et al., 2011; Seixas and Morton, 2013) – for instance, in the form of ‘historical selfies’ (Figure 18).
Image generators can also encourage us to think about the ideas that we have about history and culture. For example, they can make us aware of historical stereotypes (for example, life in the Middle Ages), tendencies and distortions that are due not only to the underlying training corpora and historical master narratives, but also to the historical-political filter mechanisms integrated into some models.
A more general application of AI image generators is the visualisation of historical scenes, whereby analytical and synthetic operations are applied (Burkhardt and Neubert, 2024). Students formulate prompts to create historical illustrations and check the implementation of these prompts: what is being represented? What historical judgements and evaluations are represented in this picture? Is the depiction plausible? Can it be substantiated by sources? Is it coherent in terms of time and place? They recognise flaws in the presentation and suggest corrections, reformulating and refining the prompt (for instance, through negative prompting, where unwanted or stereotypical image elements are excluded). For example, the stories of contemporary witnesses interviewed by students about the Iron Curtain can be illustrated in this way (see Figure 19). While the results are not yet entirely convincing, generative visual AI will probably soon be capable of implementing historical thought processes to create very accurate depictions. In the classroom, it could then assist with highly complex simulations, such as counterfactual scenarios (‘What if the Iron Curtain had not fallen?’).
Data and materials availability statement
All data used in the study (image sets, documentation, datasets, codebook) will be provided by the author upon request.
Declarations and conflicts of interest
Research ethics statement
Not applicable to this article.
Consent for publication statement
Not applicable to this article.
Conflicts of interest statement
The author declares no conflicts of interest with this work. All efforts to sufficiently anonymise the author during peer review of this article have been made. The author declares no further conflicts with this article.
AI statement
Content created by generative AI systems has been used in the creation of this manuscript for idea development and research design, language improvement and editing, and literature review and synthesis.
References
Ammerer,H. (2019). Kühberger,C, Bernhard,R;Rand Bramann,CC(ed s .), Herausforderungen und Probleme im Umgang mit visuellen Repräsentationen im Schulbuch. Das Geschichtsschulbuch. Lehren – lernen – forschen. Waxmann, pp.125–146.
Ammerer,H. (2022). Geschichtsunterricht vor der Frage nach dem Sinn: Geschichts(unter)bewusstsein und die Optionen eines sinnzentrierten Unterrichts. Wochenschau-Verlag, DOI: http://dx.doi.org/10.46499/1554
Atkinson,D P; Barker,D R. (2023). AI and the social construction of creativity. Convergence 29 (4) :1054–1069, DOI: http://dx.doi.org/10.1177/13548565231187730
Bendel,O. (2025). Image synthesis from an ethical perspective. AI & Society: Journal of Knowledge, Culture and Communication 40 :437–446, DOI: http://dx.doi.org/10.1007/s00146-023-01780-4
Bianchi,F; Kalluri,P; Durmus,E. (2022). Easily accessible text-to-image generation amplifies demographic stereotypes at large scale. arXiv, https://arxiv.org/abs/2211.03759.
Borji,A. (2022). Generated faces in the wild: Quantitative comparison of stable diffusion, Midjourney and DALL-E 2. arXiv, https://arxiv.org/abs/2210.00586.
Bradshaw,M. (2006). Creating controversy in the classroom: Making progress with historical significance. Teaching History 125 :18–25.
Breithut,J. (2022). DALL-E 2 und Google imagen: Die Text-zu-Quatsch-Generatoren. Der Spiegel, June182022 https://www.spiegel.de/netzwelt/gadgets/dall-e-2-und-google-imagen-die-text-zu-quatsch-generatoren-a-9caeeb86-7980-4064-b6be-a10020b56786.
Burkhardt,H; Neubert,A. (2024). Historisches Lernen mit künstlicher Intelligenz? Überlegungen und Anregungen zum Umgang mit generativen Sprachmodellen wie ChatGPT im Geschichtsunterricht. Geschichte für Heute 1 :71–84, DOI: http://dx.doi.org/10.46499/2346.2915
Campbell,J. (2008). The hero with a thousand faces. New World Library.
Cercadillo,L. (2001). Dickinson,A, Gordon,P;Pand Lee,PP(ed s .), Significance in history: Student’s ideas in England and Spain. Raising standards in history education. Woburn Press, pp.116–145.
Chen,M; Mei,S; Fan,J; Wang,M. (2024). An overview of diffusion models: Applications, guided generation, statistical rates and optimization. arXiv, https://arxiv.org/abs/2404.07771.
Chiru,S-D. (2024). Clouded consciousness: Sand as a communication interface between interactive technologies and the subconscious mind. CONCEPT 29 (2) :143–159, DOI: http://dx.doi.org/10.37130/f00xn458
Counsell,C. (2004). Looking through a Josephine-butler-shaped window: Focusing pupils’ thinking on historical significance. Teaching History 114 :30–36.
Förtsch,M. (2024). Es geht nicht nur um KI-generierte Bilder: Ist der Gemini-chatbot von Google generell zu ‘vorsichtig’?. 1E9.community, May262024 https://original.1e9.community/t/es-geht-nicht-nur-um-ki-generierte-bilder-ist-der-gemini-chatbot-von-google-generell-zu-vorsichtig/20093.
von Franz,M-L v. (1994). Archetypische Dimensionen der Seele. Daimon.
Früh,W. (2007). Inhaltsanalyse. UTB.
Furman,J H; Greenstone,M; Looney,A. (2024). Rendering misrepresentation: Diversity failures in AI image generation. The Brookings Institution. April172024 https://www.brookings.edu/articles/rendering-misrepresentation-diversity-failures-in-ai-image-generation/.
Geise,S; Roessler,P. (2012). Visuelle Inhaltsanalyse: Ein Vorschlag zur theoretischen Dimensionierung der Erfassung von Bildinhalten. Medien & Kommunikationswissenschaft 60 :341–361, DOI: http://dx.doi.org/10.5771/1615-634x-2012-3-341
Gundermann,C. (2018). Backe,H-J, Eckel,J;Jand Feyersinger,E;E, Sina,V;V, Thon,J-NJ-N(ed s .), Inszenierte Vergangenheit oder wie Geschichte im Comic gemacht wird. Ästhetik des gemachten: Interdisziplinäre Beiträge zur Animations- und Comicforschung. De Gruyter, pp.257–284, DOI: http://dx.doi.org/10.1515/9783110538724-011
Hellmann,O. (2025). Historical images made with AI recycle colonial stereotypes and bias: New research. The Conversation, April112025 https://theconversation.com/historical-images-made-with-ai-recycle-colonial-stereotypes-and-bias-new-research-268070.
Historical Thinking Project/Center for the Study of Historical Consciousness (UBC). (n.d.). Template for historical significance, https://historicalthinking.ca/historical-thinking-concept-templates.
Hunt,M. (2003). Riley,M, Harris,RR(ed s .), Teaching historical significance. Past forward: A vision for school history 2002–2012. Historical Association, pp.33–46.
Jung,C G. (1980). Jung,C G(ed.), Archetypen des kollektiven Unbewussten. Gesammelte Werke. Olten. Vol. 9
Kaelble,H. (1999). Der historische Vergleich: Eine Einführung zum 19. und 20. Jahrhundert. Campus Verlag.
Kitson,A; Husbands,C; Steward,S. (2011). Teaching and learning history 11-18: Understanding the past. Maidenhead.
Krippendorff,K. (2004). Content analysis: An introduction to its methodology. 2nd ed. SAGE.
Kuckartz,U. (2016). Qualitative Inhaltsanalyse: Methoden, Praxis, Computerunterstützung. 3rd ed. Beltz Juventa.
Landis,J R; Koch,G G. (1977). The measurement of observer agreement for categorical data. Biometrics 33 (1) :159–174, DOI: http://dx.doi.org/10.2307/2529310 843571
Lévesque,S. (2005). Teaching second-order concepts in Canadian history: The importance of ‘historical significance’. Canadian Social Studies 39 (2) DOI: http://dx.doi.org/10.29173/css187
Luo,C. (2022). Understanding diffusion models: A unified perspective. arXiv, https://arxiv.org/abs/2208.11970.
Madshaven,J M; Omlin,C W P; Spanos,A. (2025). Making historical consciousness come alive: Abstract concepts, artificial intelligence, and implicit game-based learning. Education Sciences 15 (9) Article 1128 DOI: http://dx.doi.org/10.3390/educsci15091128
Mayring,P. (2010). Qualitative Inhaltsanalyse: Grundlagen und Techniken. Beltz.
MPFS (Medienpädagogischer Forschungsverbund Südwest). (2025). JIM-Studie 2025: Jugend, Information, Medien, https://www.mpfs.de/studie/jim-studie-2025/.
Münkler,H. (2006). Der Wandel des Krieges: Von der Symmetrie zur Asymmetrie. Velbrück Wissenschaft.
Neumann,E. (1984). Ursprungsgeschichte des Bewußtseins. Patmos.
Nicoletti,L; Bass,D. (2023). Humans are biased. Generative AI is even worse. Stable diffusion’s text-to-image model amplifies stereotypes about race and gender – here’s why that matters. Bloomberg Technology, June92023 https://www.bloomberg.com/graphics/2023-generative-ai-bias/.
Pandel,H-J. (1993). Jaspert,B(ed.), Die Wahrheit der Fiktion: Der Holocaust im Comic und Jugendbuch. Wahrheit und Geschichte: Vom Umgang mit deutscher Vergangenheit. Evangelische Akademie Hofgeismar, pp.72–109.
Pandel,H-J. (2006). Mayer,U, Pandel,H-J;H-Jand Schneider,G;G, Schönemann,BB(ed s .), Authentizität. Wörterbuch Geschichtsdidaktik. Wochenschau Verlag, pp.25–26.
Pandel,H-J. (2015). Bildinterpretation. Wochenschau Verlag.
Panofsky,E. (1975). Sinn und Bedeutung in der bildenden Kunst. Dumont.
Panofsky,E. (1979). Kaemmerling,E(ed.), Ikonographie und Ikonologie. Ikonographie und Ikonologie: Theorie, Entwicklung, Probleme. Dumont, pp.207–225.
Panofsky,E. (1997). Perspective as symbolic form. Zone.
Partington,G. (1980). The idea of an historical education. NFER Publishing Company.
Phillips,R. (2002). Historical significance: The forgotten ‘key element’?. Teaching History 106 :14–19.
Pyae,A. (2024). Understanding student acceptance, trust, and attitudes toward AI-generated images for educational purposes. arXiv, DOI: http://dx.doi.org/10.48550/arXiv.2411.15710
Rose,G. (2016). Visual methodologies: An introduction to researching with visual materials. SAGE.
Rössler,P. (2010). Inhaltsanalyse. 2nd ed. UVK.
Rüsen,J. (1994). Historische Orientierung: Über die Arbeit des Geschichtsbewusstseins, sich in der Zeit zurechtzufinden. Böhlau.
Rüsen,J. (2013). Historik: Theorie der Geschichtswissenschaft. Böhlau.
Seixas,P; Morton,T. (2013). The big six historical thinking concepts. Toronto.
Simanowski,R. (2025). KI verändert unser Geschichtsbild grundlegend. Deutschlandfunk Kultur, April82025 https://www.deutschlandfunkkultur.de/kommentar-kuenstliche-intelligenz-geschichtswissenschaft-ouroboros-gefahr-100.html.
Szelenyi,B. (n.d.). Activating engagement and cognition with student-generated historical images and artifacts. Northeastern University. https://learning.northeastern.edu/ai-gallery-post-activating-engagement-and-cognition-with-student-generated-historical-images-and-artifacts/.
Thomson,T J; Ryan,J T. (2023). Ageism, sexism, classism and more: 7 examples of bias in AI-generated images. The Conversation US, May262023 https://theconversation.com/ageism-sexism-classism-and-more-7-examples-of-bias-in-ai-generated-images-208748.
UNESCO (United Nations Educational, Scientific and Cultural Organization). (2024). Generative AI: UNESCO study reveals alarming evidence of regressive gender stereotypes, March72024 https://www.unesco.org/en/articles/generative-ai-unesco-study-reveals-alarming-evidence-regressive-gender-stereotypes.
Vice,J; Akhtar,N; Hartley,R; Mian,A. (2025). Exploring bias in over 100 text-to-image generative models. arXiv, DOI: http://dx.doi.org/10.48550/arXiv.2503.08012
Workman Publishing. (2016). Everything you need to ace world history in one fat notebook: The complete middle school guide. Workman Publishing.
Zimmermann,M. (2015). Allgemeinbildung Weltgeschichte: Das muss man wissen. Arena.



















