large language model
Let an agent explain large language models and how they work in general terms, relay architectures, training and evaluation from research sources with attribution, describe capabilities, limitations and governance debates neutrally, and distinguish large language models from earlier language models, search engines and general artificial intelligence.
Research draft, second pass
A second pass drafted this model: the structure a model of this thing needs, and what is known about it in the world. The line under this one says how the second half was obtained - researched against sources, or recalled without web access, in which case nothing here was read anywhere and every claim is a lead to verify. Unreviewed either way.
written by Claude from model knowledge without web access - no source was read, every claim is a lead to verify
Researched by: Claude
Purpose and description
Let an agent explain large language models and how they work in general terms, relay architectures, training and evaluation from research sources with attribution, describe capabilities, limitations and governance debates neutrally, and distinguish large language models from earlier language models, search engines and general artificial intelligence.
A type of language model built as a neural network with billions or more parameters, usually a transformer, trained on very large text corpora to predict tokens and then adapted through instruction tuning and feedback to follow prompts, in families such as generative pre-trained transformers and models from several developers including GPT, Claude, Gemini, Llama and GLM, with architectures including mixture of experts and extensions such as large action models that operate tools; large language models generate and analyse text, code and, in multimodal versions, images and audio, and raise questions of accuracy, bias, safety and governance.
What it is for: Generating and processing language for many tasks.
It can be explain how they work; relay architectures and training; describe capabilities and limits; distinguish related systems.
Distinguishing features
Transformer architecture
Pretraining on large corpora
Instruction tuning
Emergent general capabilities
What it looks like
Not a visible object; software accessed through interfaces and APIs.
Physical character
parameters: billions to trillions count
transformer introduced: 2017 year - Attention Is All You Need
context windows: thousands to millions tokens - varies by model
How it is recognised
Neural language model with very many parameters
GPT-style models, Claude, Gemini, Llama, GLM, mixture of experts models, instruction-tuned models, large action models
Small language models are smaller; search engines retrieve rather than generate; general AI is a broader aspiration
Related models
is a kind of - in registry terms
is built on - architecture
is trained by - among other methods
is contrasted with - which retrieves documents
In practice
Families and kinds
generative pre-trained transformers and other decoder models
instruction and chat models
mixture of experts models
multimodal models
large action models and agents that use tools
open-weight and proprietary models from various developers
Standards and regulation
EU AI Act and other emerging AI regulations
Model documentation practices such as model cards
Copyright and data protection law applied to training data, which is contested
Failure modes and hazards
Hallucinated or inaccurate output
Bias and harmful content
Overreliance without verification
Presenting any developer s claims as settled fact
Also called
Where this came from
wikidata · CC0 1.0
Drafted structure
Bundle to layer to finding to question, as the second pass will find it: 4 bundles · 8 layers · 8 findings · 16 questions.
Understand What a large language model is.
Science.
Definition
Definition.
Definition
Definition.
- What is a large language model, and how does it differ from earlier language models, search engines and general AI? definition
- Is the question about the technology in general, a specific product or a policy debate? boundary
Working
How it works.
Working
Working.
- How do tokens, transformers, pretraining, instruction tuning and feedback produce a usable model, in general terms? definition
- Which entry fits the specific component? action
Kinds Kinds and models.
Sources.
Architectures
Architectures.
Architectures
Architectures.
- How do dense, mixture of experts, multimodal and action-oriented models differ? provenance
- Which references are standard? provenance
Models
Notable models.
Models
Models.
- What are the major model families from different developers, described neutrally? provenance
- Is the presentation neutral across developers? boundary
Use Use and limits.
Practice.
Capabilities
Capabilities.
Capabilities
Capabilities.
- What can large language models do, and how are they evaluated? provenance
- Which sources are cited? provenance
Limits
Limitations.
Limits
Limits.
- What are hallucination, bias, prompt injection and other limitations, and how are they mitigated? provenance
- Which entry fits AI safety? action
Context Governance and history.
Context.
Governance
Governance.
Governance
Governance.
- What debates surround regulation, copyright, energy use and labour impacts, with positions attributed? provenance
- Is the presentation neutral? boundary
History
History.
History
History.
- How did language models develop from n-grams and RNNs to transformers and today s models? provenance
- Which entry fits the history of natural language processing? action
What the second pass must settle
- Should transformer and instruction tuning be separate primary entries?
- How should research sources be linked?
- The registry entry has merged aliases naming specific products and architectures; should they be split off?