Skip to main content

What is Generative AI?

Attribution

Original work: "Educators' guide to multimodal learning and Generative AI" — Tünde Varga-Atkins, Samuel Saunders, et al. (2024/25) — CC BY-NC 4.0
Adapted for UK Nursing Education by: Lincoln Gombedza, RN (LD)

Before proceeding, it is useful to explain what we mean by Generative AI. In fact, this is a useful question to discuss with colleagues and students before using it any way, as there are many common misconceptions.

Common Misconceptions​

GenAI is NOT a Search Engine​

Many assistants can now search the web and show citations, and search engines now put AI-written summaries above their results. That makes the difference harder to see, but it still matters.

FeatureGoing to the source (e.g., NICE, BNF)Generative AI (including AI search summaries)
Primary GoalPublish accurate, accountable informationGenerate a plausible, relevant response
Source Verification✅ The source is the evidence⚠️ May cite sources, but citations can be wrong or not support the claim
Accuracy for Clinical Data✅ Reviewed and dated⚠️ Can hallucinate or miss context
Best ForChecking guidance before you act or teachBrainstorming, drafting, summarising, explaining
Critical Reminder

GenAI is unreliable as a way of finding clinical information, because it generates a plausible answer rather than retrieving a checked one. In January 2026 Google removed AI Overviews from some health searches after a Guardian investigation found inaccurate health advice, including on cancer screening and liver blood tests (TechCrunch).

For nursing educators: This is particularly important when students might use GenAI to look up clinical information or evidence-based guidelines. Always emphasise the need to verify information against authoritative sources like:

  • NICE guidelines
  • NMC standards
  • Cochrane reviews
  • Peer-reviewed nursing journals

How GenAI Actually Works​

GenAI employs deep machine learning techniques to process information contained within huge datasets to generate outputs based on human prompts inputted by the user.

The "Next Token Prediction" Model​

Most GenAI we refer to are Large Language Models (LLMs) trained on vast amounts of text data to decode, generate, and manipulate human language. Current "reasoning" models work through a problem step by step before answering, which improves many answers but doesn't stop them being confidently wrong. GenAI is also increasingly capable of producing multimodal content.


Multimodal Capabilities​

GenAI can create content across multiple formats:

ModeExamplesNursing Use Case
📝 TextEssays, care plansPatient documentation drafts
🗣️ SpeechText-to-speechAccessibility for learning disabilities
🎵 AudioPodcasts, narrationsCommute-friendly lectures
🖼️ ImagesDiagrams, infographicsAnatomy illustrations
🎥 VideoDemonstrationsClinical skill tutorials
🎭 3D ModelsAnatomical structuresInteractive learning

Anatomy of a Prompt​

To get the best results from these models, you need to structure your requests effectively. Hover over the parts of this prompt to understand their function:

"Act as aSenior Nurse, explainasthma managementto anewly diagnosed teenagerusingsupportive bullet points."
Hover over the coloured words above to see what they do.

GenAI technology enables users to create, manipulate, and adapt content and integrate different semiotic forms to produce multimodal artefacts, and thus can be embedded into pedagogical practices that already emphasise diverse modes of engagement.


Defining GenAI's Role​

GenAI's rapid development has been accompanied by suggestions on how to define and use this technology:

Co-Intelligence (Mollick, 2024)​

Mollick, E. (2024) Co-Intelligence: Living and Working with AI. London: WH Allen. Suggests we should consider GenAI to be a 'co-intelligence' that works alongside human intelligence.

Co-Creator (Cope & Kalantzis, 2024)​

Cope, B. and Kalantzis, M. (2024) 'A multimodal grammar of artificial intelligence: measuring the gains and losses in generative AI', Multimodality & Society, 4(2), pp. 123–152. doi:10.1177/26349795231221699 Suggests that we should understand it as a unique 'co-creator' that works alongside users in an assistive but unique role. They call this cyber-social learning – a collaborative partnership between human and machine intelligences, each with distinct, but complementary, strengths for completing an activity.

They suggest that this collaboration 'enables new processes for knowledge creation', where educators and students learn by evaluating, refining and re-imagining AI outputs, assembling them into multimodal artefacts.


The Human-AI Partnership​

Critical Perspective

While human and artificial intelligences can work together and have a unique role to play within a cyber-social partnership, the term 'intelligence', taken as part of 'Generative AI' can, and perhaps should, be challenged in favour of more specific computer-science-driven terminologies, such as LLMs (large language models).

'Intelligence' implies 'consciousness' that AI simply does not have, despite its (and their parent companies') attempts to lead users into believing it does.

It is GenAI's very lack of 'intelligence', either emotional or intellectual, that highlights how it operates in the cyber-social relationship with a human, whereby both parties occupy unique but symbiotic positions and consequently complement each other.

What GenAI Can Do:​

  • ✅ Generate and transform multimodal content at scale
  • ✅ Work with extreme efficiency
  • ✅ Process large amounts of information quickly
  • ✅ Create drafts, artefacts and prototypes

What GenAI Cannot Do (Reliably, or At All):​

  • ⚠️ Grasp the full context of a real patient or situation; it only sees what it is told
  • ⚠️ Recognise emotion reliably; it can describe feelings in text, but can't read the room
  • ❌ Know when it is wrong; it can sound equally confident either way
  • ❌ Make ethical judgements it can be held to account for
  • ❌ Provide clinical judgement or professional accountability

The Bottom Line: GenAI can offer raw material – drafts, artefacts and prototypes - but educators and learners are needed to bring vision, purpose, nuanced critique and meaning-making.


Types of Multimodal GenAI​

The guide's aim is to encapsulate strategies for the effective incorporation of GenAI in multimodal teaching, learning, and assessment and to position Generative AI most effectively within the cyber-social relationship with human users.

We interpret 'multimodal GenAI' in different ways:

1. Platform Capabilities​

Multimodal GenAI can refer to platform capabilities that utilise modalities beyond text-to-text:

Current models​

The table below comes from the toolkit's model registry, so it is updated in one place whenever providers release new models. Every major assistant is now multimodal: they read images, documents and charts as well as text.

ProviderModelBest forWhere to get itNotes for UK educators
AnthropicClaude Opus 5.5Most capableClaude.ai paid plans, Claude API, Amazon Bedrock, Google Vertex AI; rolling out in Microsoft 365 CopilotTeam and Enterprise plans do not train on your data by default; check your organisation's agreement before entering any identifiable information.
AnthropicClaude Fable 5.1Frontier (premium)Claude.ai, Claude API, AWS, Google Cloud, Microsoft Azure; Microsoft 365 Copilot Cowork and Copilot StudioAnthropic's premium general-release model; Opus 5.5 gives similar results on most tasks at lower cost.
AnthropicClaude Sonnet 5Everyday workDefault model on Claude.ai Free and Pro; Claude API and cloud platformsSonnet 5.5 has been announced for the coming weeks.
AnthropicClaude Haiku 4.5Fast and low-costClaude API and cloud platformsLow-cost model for high-volume tasks; Haiku 5.5 has been announced.
OpenAIGPT-6 AstraMost capableChatGPT Plus, Pro, Business and Enterprise; OpenAI API; AWS; Microsoft 365 Copilot Cowork and Copilot StudioIts most advanced cyber-security capabilities are restricted to a vetted access programme.
OpenAIGPT-6 SolEveryday workChatGPT paid plans, OpenAI API; rolling out in Microsoft 365 Copilot
OpenAIGPT-6 LunaFast and low-costOpenAI API; ChatGPT Free and Go users in the desktop app only
OpenAIGPT-5.6 LunaWorkplace assistantDefault model for ChatGPT Free and Go users in the chat windowMost students using free ChatGPT meet this model, not GPT-6.
GoogleGemini 3.1 ProMost capableGemini app (Google AI Pro and Ultra), Google AI Studio, Vertex AI, Google WorkspaceGoogle's newest Pro model; there was no Gemini 3.5 Pro. Gemini 4 Pro is expected but not yet released.
GoogleGemini 3.8 FlashEveryday workGemini app with Google AI Pro or Ultra; free in Google AI Studio; Gemini APIEligible students in 140+ countries, including the UK, can claim Google AI Plus free until 31 December 2026.
GoogleGemini 3.5 Flash-LiteFast and low-costGemini API, Vertex AI
GoogleGemini 3 Deep ThinkDeep reasoningGemini app with Google AI Ultra
MicrosoftMicrosoft 365 CopilotWorkplace assistantMicrosoft 365 work and education accounts, including many NHS and university tenantsUses several models: GPT-5.6 is the preferred model, with Claude Opus 5, Fable 5.1, Opus 5.5 and GPT-6 Sol and Astra in some features. Which ones you can use depends on your organisation's admin settings.
Mistral AIMistral Large 3Most capableLe Chat, Mistral API, Amazon Bedrock; open weights (Apache 2.0)EU-based provider, often raised in data-residency discussions.
Models checked 24 September 2026. They change often, so check the provider before relying on a specific version.
Image Generation
  • ChatGPT (GPT Image 2.5) and Gemini - image generation built into the assistants
  • Midjourney - artistic and photorealistic outputs
  • Adobe Firefly - trained on licensed content, suited to commercial use
  • Stable Diffusion - open-weights models you can run locally

Never generate images that could be mistaken for real patients, and check clinical images (wounds, rashes, anatomy) against authoritative sources: AI images often get details wrong.

Text-to-Speech
  • ElevenLabs - realistic voices, including voice cloning
  • Google Cloud Text-to-Speech - natural-sounding, many languages
  • Azure AI Speech - enterprise speech services

Only clone a voice with that person's explicit consent.

Text-to-Video
  • Google Veo and OpenAI Sora - high-fidelity video from text or images
  • Runway - creative video generation and editing
  • Synthesia and HeyGen - AI avatar presenters
Speech-to-Text
  • Whisper (OpenAI) - open-source transcription
  • Microsoft Teams / Copilot and Otter.ai - meeting transcripts and summaries
  • Ambient voice technology ("AI scribes") - clinical note drafting; see AI scribes on placement
Image-to-Text & Vision
  • ChatGPT (GPT-6 Astra, GPT-5.6 Luna) - Understanding images, charts and documents
  • Gemini (Gemini 3.1 Pro) - Native multimodal: text, images, audio and video
  • Claude (Claude Opus 5.5) - Document, chart and image analysis
Rapidly Evolving Field

AI models evolve extremely rapidly. The model table above shows the date it was last checked. Always check the latest releases from:

Nursing Example: Creating visual care pathways from text descriptions, or transcribing verbal patient handovers.

2. Multimodal Learning Activities​

A multimodal learning or teaching activity itself (e.g. a lecture or a virtual simulation) that utilises GenAI within its process (whether GenAI itself is text-to-text or multimodal).

Nursing Example: A simulation where students interact with an AI-generated patient avatar that responds via text and speech.

3. Modal Conversion​

Using GenAI to convert one artefact/modality (e.g. slides or images) into another modality (e.g. text or sound).

Nursing Example: Converting a PowerPoint lecture on wound care into a podcast for students to listen to during commute.


The Multimodality Continuum​

The following illustrates GenAI capabilities' development in terms of educators' uses of multimodal resources from text to immersive simulation:

GenAI now offers real-time interaction with avatars, voices and personas, and agents that carry out multi-step tasks. Only a couple of years ago, most tools produced static text, audio or images.

Rapid Evolution

This guide offers broad principles and approaches rather than platform-specific suggestions to ensure relevance across disciplines and learning contexts. Even during the lifespan of this project (2024/25), GenAI's multimodal capabilities have evolved so rapidly that listing very specific concrete examples risks the information becoming quickly outdated.


For Nursing Educators: Key Takeaways​

  1. GenAI is not a search engine — Don't let students use it as one for clinical information
  2. GenAI lacks clinical judgement — Human review and verification is essential
  3. GenAI is a tool, not a replacement — It augments, not replaces, nursing expertise
  4. Multimodal possibilities are vast — From text to video to simulations
  5. Evolution is rapid — Stay current but focus on principles, not specific platforms

Next: Continue to the Main Introduction for a comprehensive overview of multimodal learning and nursing context.