What is Generative AI?
Original work: "Educators' guide to multimodal learning and Generative AI" — Tünde Varga-Atkins, Samuel Saunders, et al. (2024/25) — CC BY-NC 4.0
Adapted for UK Nursing Education by: Lincoln Gombedza, RN (LD)
Before proceeding, it is useful to explain what we mean by Generative AI. In fact, this is a useful question to discuss with colleagues and students before using it any way, as there are many common misconceptions.
Common Misconceptions
GenAI is NOT a Search Engine
Many assistants can now search the web and show citations, and search engines now put AI-written summaries above their results. That makes the difference harder to see, but it still matters.
| Feature | Going to the source (e.g., NICE, BNF) | Generative AI (including AI search summaries) |
|---|---|---|
| Primary Goal | Publish accurate, accountable information | Generate a plausible, relevant response |
| Source Verification | ✅ The source is the evidence | ⚠️ May cite sources, but citations can be wrong or not support the claim |
| Accuracy for Clinical Data | ✅ Reviewed and dated | ⚠️ Can hallucinate or miss context |
| Best For | Checking guidance before you act or teach | Brainstorming, drafting, summarising, explaining |
GenAI is unreliable as a way of finding clinical information, because it generates a plausible answer rather than retrieving a checked one. In January 2026 Google removed AI Overviews from some health searches after a Guardian investigation found inaccurate health advice, including on cancer screening and liver blood tests (TechCrunch).
For nursing educators: This is particularly important when students might use GenAI to look up clinical information or evidence-based guidelines. Always emphasise the need to verify information against authoritative sources like:
- NICE guidelines
- NMC standards
- Cochrane reviews
- Peer-reviewed nursing journals
How GenAI Actually Works
GenAI employs deep machine learning techniques to process information contained within huge datasets to generate outputs based on human prompts inputted by the user.
The "Next Token Prediction" Model
Most GenAI we refer to are Large Language Models (LLMs) trained on vast amounts of text data to decode, generate, and manipulate human language. Current "reasoning" models work through a problem step by step before answering, which improves many answers but doesn't stop them being confidently wrong. GenAI is also increasingly capable of producing multimodal content.
Multimodal Capabilities
GenAI can create content across multiple formats:
| Mode | Examples | Nursing Use Case |
|---|---|---|
| 📝 Text | Essays, care plans | Patient documentation drafts |
| 🗣️ Speech | Text-to-speech | Accessibility for learning disabilities |
| 🎵 Audio | Podcasts, narrations | Commute-friendly lectures |
| 🖼️ Images | Diagrams, infographics | Anatomy illustrations |
| 🎥 Video | Demonstrations | Clinical skill tutorials |
| 🎭 3D Models | Anatomical structures | Interactive learning |
Anatomy of a Prompt
To get the best results from these models, you need to structure your requests effectively. Hover over the parts of this prompt to understand their function:
GenAI technology enables users to create, manipulate, and adapt content and integrate different semiotic forms to produce multimodal artefacts, and thus can be embedded into pedagogical practices that already emphasise diverse modes of engagement.
Defining GenAI's Role
GenAI's rapid development has been accompanied by suggestions on how to define and use this technology:
Co-Intelligence (Mollick, 2024)
Mollick, E. (2024) Co-Intelligence: Living and Working with AI. London: WH Allen. Suggests we should consider GenAI to be a 'co-intelligence' that works alongside human intelligence.
Co-Creator (Cope & Kalantzis, 2024)
Cope, B. and Kalantzis, M. (2024) 'A multimodal grammar of artificial intelligence: measuring the gains and losses in generative AI', Multimodality & Society, 4(2), pp. 123–152. doi:10.1177/26349795231221699 Suggests that we should understand it as a unique 'co-creator' that works alongside users in an assistive but unique role. They call this cyber-social learning – a collaborative partnership between human and machine intelligences, each with distinct, but complementary, strengths for completing an activity.
They suggest that this collaboration 'enables new processes for knowledge creation', where educators and students learn by evaluating, refining and re-imagining AI outputs, assembling them into multimodal artefacts.
The Human-AI Partnership
While human and artificial intelligences can work together and have a unique role to play within a cyber-social partnership, the term 'intelligence', taken as part of 'Generative AI' can, and perhaps should, be challenged in favour of more specific computer-science-driven terminologies, such as LLMs (large language models).
'Intelligence' implies 'consciousness' that AI simply does not have, despite its (and their parent companies') attempts to lead users into believing it does.
It is GenAI's very lack of 'intelligence', either emotional or intellectual, that highlights how it operates in the cyber-social relationship with a human, whereby both parties occupy unique but symbiotic positions and consequently complement each other.
What GenAI Can Do:
- ✅ Generate and transform multimodal content at scale
- ✅ Work with extreme efficiency
- ✅ Process large amounts of information quickly
- ✅ Create drafts, artefacts and prototypes
What GenAI Cannot Do (Reliably, or At All):
- ⚠️ Grasp the full context of a real patient or situation; it only sees what it is told
- ⚠️ Recognise emotion reliably; it can describe feelings in text, but can't read the room
- ❌ Know when it is wrong; it can sound equally confident either way
- ❌ Make ethical judgements it can be held to account for
- ❌ Provide clinical judgement or professional accountability
The Bottom Line: GenAI can offer raw material – drafts, artefacts and prototypes - but educators and learners are needed to bring vision, purpose, nuanced critique and meaning-making.
Types of Multimodal GenAI
The guide's aim is to encapsulate strategies for the effective incorporation of GenAI in multimodal teaching, learning, and assessment and to position Generative AI most effectively within the cyber-social relationship with human users.
We interpret 'multimodal GenAI' in different ways:
1. Platform Capabilities
Multimodal GenAI can refer to platform capabilities that utilise modalities beyond text-to-text:
Current models
The table below comes from the toolkit's model registry, so it is updated in one place whenever providers release new models. Every major assistant is now multimodal: they read images, documents and charts as well as text.
| Provider | Model | Best for | Where to get it | Notes for UK educators |
|---|---|---|---|---|
| Anthropic | Claude Opus 5.5 | Most capable | Claude.ai paid plans, Claude API, Amazon Bedrock, Google Vertex AI; rolling out in Microsoft 365 Copilot | Team and Enterprise plans do not train on your data by default; check your organisation's agreement before entering any identifiable information. |
| Anthropic | Claude Fable 5.1 | Frontier (premium) | Claude.ai, Claude API, AWS, Google Cloud, Microsoft Azure; Microsoft 365 Copilot Cowork and Copilot Studio | Anthropic's premium general-release model; Opus 5.5 gives similar results on most tasks at lower cost. |
| Anthropic | Claude Sonnet 5 | Everyday work | Default model on Claude.ai Free and Pro; Claude API and cloud platforms | Sonnet 5.5 has been announced for the coming weeks. |
| Anthropic | Claude Haiku 4.5 | Fast and low-cost | Claude API and cloud platforms | Low-cost model for high-volume tasks; Haiku 5.5 has been announced. |
| OpenAI | GPT-6 Astra | Most capable | ChatGPT Plus, Pro, Business and Enterprise; OpenAI API; AWS; Microsoft 365 Copilot Cowork and Copilot Studio | Its most advanced cyber-security capabilities are restricted to a vetted access programme. |
| OpenAI | GPT-6 Sol | Everyday work | ChatGPT paid plans, OpenAI API; rolling out in Microsoft 365 Copilot | |
| OpenAI | GPT-6 Luna | Fast and low-cost | OpenAI API; ChatGPT Free and Go users in the desktop app only | |
| OpenAI | GPT-5.6 Luna | Workplace assistant | Default model for ChatGPT Free and Go users in the chat window | Most students using free ChatGPT meet this model, not GPT-6. |
| Gemini 3.1 Pro | Most capable | Gemini app (Google AI Pro and Ultra), Google AI Studio, Vertex AI, Google Workspace | Google's newest Pro model; there was no Gemini 3.5 Pro. Gemini 4 Pro is expected but not yet released. | |
| Gemini 3.8 Flash | Everyday work | Gemini app with Google AI Pro or Ultra; free in Google AI Studio; Gemini API | Eligible students in 140+ countries, including the UK, can claim Google AI Plus free until 31 December 2026. | |
| Gemini 3.5 Flash-Lite | Fast and low-cost | Gemini API, Vertex AI | ||
| Gemini 3 Deep Think | Deep reasoning | Gemini app with Google AI Ultra | ||
| Microsoft | Microsoft 365 Copilot | Workplace assistant | Microsoft 365 work and education accounts, including many NHS and university tenants | Uses several models: GPT-5.6 is the preferred model, with Claude Opus 5, Fable 5.1, Opus 5.5 and GPT-6 Sol and Astra in some features. Which ones you can use depends on your organisation's admin settings. |
| Mistral AI | Mistral Large 3 | Most capable | Le Chat, Mistral API, Amazon Bedrock; open weights (Apache 2.0) | EU-based provider, often raised in data-residency discussions. |
Image Generation
- ChatGPT (GPT Image 2.5) and Gemini - image generation built into the assistants
- Midjourney - artistic and photorealistic outputs
- Adobe Firefly - trained on licensed content, suited to commercial use
- Stable Diffusion - open-weights models you can run locally
Never generate images that could be mistaken for real patients, and check clinical images (wounds, rashes, anatomy) against authoritative sources: AI images often get details wrong.
Text-to-Speech
- ElevenLabs - realistic voices, including voice cloning
- Google Cloud Text-to-Speech - natural-sounding, many languages
- Azure AI Speech - enterprise speech services
Only clone a voice with that person's explicit consent.
Text-to-Video
- Google Veo and OpenAI Sora - high-fidelity video from text or images
- Runway - creative video generation and editing
- Synthesia and HeyGen - AI avatar presenters
Speech-to-Text
- Whisper (OpenAI) - open-source transcription
- Microsoft Teams / Copilot and Otter.ai - meeting transcripts and summaries
- Ambient voice technology ("AI scribes") - clinical note drafting; see AI scribes on placement
Image-to-Text & Vision
- ChatGPT (GPT-6 Astra, GPT-5.6 Luna) - Understanding images, charts and documents
- Gemini (Gemini 3.1 Pro) - Native multimodal: text, images, audio and video
- Claude (Claude Opus 5.5) - Document, chart and image analysis
AI models evolve extremely rapidly. The model table above shows the date it was last checked. Always check the latest releases from:
Nursing Example: Creating visual care pathways from text descriptions, or transcribing verbal patient handovers.
2. Multimodal Learning Activities
A multimodal learning or teaching activity itself (e.g. a lecture or a virtual simulation) that utilises GenAI within its process (whether GenAI itself is text-to-text or multimodal).
Nursing Example: A simulation where students interact with an AI-generated patient avatar that responds via text and speech.
3. Modal Conversion
Using GenAI to convert one artefact/modality (e.g. slides or images) into another modality (e.g. text or sound).
Nursing Example: Converting a PowerPoint lecture on wound care into a podcast for students to listen to during commute.
The Multimodality Continuum
The following illustrates GenAI capabilities' development in terms of educators' uses of multimodal resources from text to immersive simulation:
GenAI now offers real-time interaction with avatars, voices and personas, and agents that carry out multi-step tasks. Only a couple of years ago, most tools produced static text, audio or images.
This guide offers broad principles and approaches rather than platform-specific suggestions to ensure relevance across disciplines and learning contexts. Even during the lifespan of this project (2024/25), GenAI's multimodal capabilities have evolved so rapidly that listing very specific concrete examples risks the information becoming quickly outdated.
For Nursing Educators: Key Takeaways
- GenAI is not a search engine — Don't let students use it as one for clinical information
- GenAI lacks clinical judgement — Human review and verification is essential
- GenAI is a tool, not a replacement — It augments, not replaces, nursing expertise
- Multimodal possibilities are vast — From text to video to simulations
- Evolution is rapid — Stay current but focus on principles, not specific platforms
Next: Continue to the Main Introduction for a comprehensive overview of multimodal learning and nursing context.