Updated: Mar 23, 2026 | Read time 7 min

GPT-4: How Does It Work? Architecture, Training & Capabilities Explained

GPT-4 is a multimodal large language model from OpenAI that accepts both text and images as input. This technical deep-dive explains how GPT-4 works — covering its transformer architecture, RLHF training, and visual reasoning capabilities.
GPT4 - What is it?
John Hughes
John HughesAccuracy Team Lead
Lawrence Atkins
Lawrence AtkinsMachine Learning Engineer
Footnotes * The GPT-4 technical report contains no detail about "architecture (including model size), hardware, training compute, dataset construction, [or] training method" due to the "competitive landscape and the safety implications" of large LMs.

** A token can be either a word or sub-word, generated using a byte-pair encoding (BPE). BPE creates a smaller vocabulary by breaking down words into smaller sub-word units. It iteratively merges the most frequent pair of consecutive bytes in a corpus of text, gradually reducing the number of unique byte sequences until a desired vocabulary size is reached.

References [1] Brown, Tom, et al. "Language models are few-shot learners." Advances in neural information processing systems 33 (2020): 1877-1901.

[2] Huang, Shaohan, et al. "Language Is Not All You Need: Aligning Perception with Language Models." arXiv preprint arXiv:2302.14045 (2023).

[3] Alayrac, Jean-Baptiste, et al. "Flamingo: a visual language model for few-shot learning." arXiv preprint arXiv:2204.14198 (2022).

[4] Ramesh, Aditya, et al. "Hierarchical text-conditional image generation with clip latents." arXiv preprint arXiv:2204.06125 (2022).

[5] Radford, Alec, et al. "Robust speech recognition via large-scale weak supervision." arXiv preprint arXiv:2212.04356 (2022).

[6] Vaswani, Ashish, et al. "Attention is all you need." Advances in neural information processing systems 30 (2017).

[7] Hao, Yaru, et al. "Language models are general-purpose interfaces." arXiv preprint arXiv:2206.06336 (2022).

[8] Dosovitskiy, Alexey, et al. "An image is worth 16x16 words: Transformers for image recognition at scale." arXiv preprint arXiv:2010.11929 (2020).

[9] Radford, Alec, et al. "Learning transferable visual models from natural language supervision." International conference on machine learning. PMLR, 2021.

[10] Ouyang, Long, et al. "Training language models to follow instructions with human feedback." arXiv preprint arXiv:2203.02155 (2022).

[11] Bai, Yuntao, et al. "Constitutional AI: Harmlessness from AI Feedback." arXiv preprint arXiv:2212.08073 (2022).

[12] OpenAI. "GPT-4 Technical Report", OpenAI (2023)

[13] Wei, Jason, et al. "Chain of thought prompting elicits reasoning in large language models." arXiv preprint arXiv:2201.11903 (2022).
Author John Hughes & Lawrence Atkins
Acknowledgements Edward Rees, Ellena Reid, Liam Steadman, Markus Hennerbichler
Carousel slide image
Use Cases

Tresic selects Speechmatics to power the speech layer of its conversation intelligence platform

Speechmatics
SpeechmaticsEditorial Team
[alt: Illustration representing multilingual code-switching for the Speechmatics Melia 1 speech-to-text model.]
Technical

Melia 1 leads speech-to-text code-switching in Arabic, Mandarin and Tamil

Speechmatics
SpeechmaticsEditorial Team
[alt: Dark grid background with a circular symbol on the left and a pixelated "K" on the right connected by a cyan line.]
Product

Speechmatics & LiveKit Inference target the accuracy gap that breaks voice agents in production

Speechmatics
SpeechmaticsEditorial Team
Carousel slide image
Product

Speechmatics on Zapier: No-Code Speech-to-Text Automation

Speechmatics
SpeechmaticsEditorial Team
[alt: Medical model header asset]
Product

Speechmatics launches Medical Model for real-time clinical transcription

Speechmatics
SpeechmaticsEditorial team
Man and woman wearing headsets and talking over the phone while working on laptops.
General

AI Voice Agents vs Traditional IVR Systems: Which Is Right for Your Business?

Speechmatics
SpeechmaticsEditorial team