← Back to catalogue
Research draft

large language model

vr.tr.large-language-model · INF.KNW

Let an agent explain large language models and how they work in general terms, relay architectures, training and evaluation from research sources with attribution, describe capabilities, limitations and governance debates neutrally, and distinguish large language models from earlier language models, search engines and general artificial intelligence.

Thing Registry Information and virtual systems

Research draft, second pass

A second pass drafted this model: the structure a model of this thing needs, and what is known about it in the world. The line under this one says how the second half was obtained - researched against sources, or recalled without web access, in which case nothing here was read anywhere and every claim is a lead to verify. Unreviewed either way.

written by Claude from model knowledge without web access - no source was read, every claim is a lead to verify

Researched by: Claude

Purpose and description

Let an agent explain large language models and how they work in general terms, relay architectures, training and evaluation from research sources with attribution, describe capabilities, limitations and governance debates neutrally, and distinguish large language models from earlier language models, search engines and general artificial intelligence.

A type of language model built as a neural network with billions or more parameters, usually a transformer, trained on very large text corpora to predict tokens and then adapted through instruction tuning and feedback to follow prompts, in families such as generative pre-trained transformers and models from several developers including GPT, Claude, Gemini, Llama and GLM, with architectures including mixture of experts and extensions such as large action models that operate tools; large language models generate and analyse text, code and, in multimodal versions, images and audio, and raise questions of accuracy, bias, safety and governance.

What it is for: Generating and processing language for many tasks.

It can be explain how they work; relay architectures and training; describe capabilities and limits; distinguish related systems.

Distinguishing features

Transformer architecture

Pretraining on large corpora

Instruction tuning

Emergent general capabilities

What it looks like

Not a visible object; software accessed through interfaces and APIs.

Physical character

parameters: billions to trillions count

transformer introduced: 2017 year - Attention Is All You Need

context windows: thousands to millions tokens - varies by model

How it is recognised

Neural language model with very many parameters

GPT-style models, Claude, Gemini, Llama, GLM, mixture of experts models, instruction-tuned models, large action models

Small language models are smaller; search engines retrieve rather than generate; general AI is a broader aspiration

Related models

is a kind of - in registry terms

language model

is built on - architecture

transformer (machine learning)

is trained by - among other methods

reinforcement learning from human feedback

is contrasted with - which retrieves documents

search engine

In practice

Families and kinds

generative pre-trained transformers and other decoder models

instruction and chat models

mixture of experts models

multimodal models

large action models and agents that use tools

open-weight and proprietary models from various developers

Standards and regulation

EU AI Act and other emerging AI regulations

Model documentation practices such as model cards

Copyright and data protection law applied to training data, which is contested

Failure modes and hazards

Hallucinated or inaccurate output

Bias and harmful content

Overreliance without verification

Presenting any developer s claims as settled fact

Also called

large action modelGLMmixture of experts modelgenerative pre-trained transformerClaudeinstruction modelGeminifamily of large language modelsmultimodal large language modelreasoning language model1.58-bit large language modelMolmoDNA large language modelcustom GPTClaude Mythos PreviewClaude Mythos

Where this came from

wikidata · CC0 1.0

Drafted structure

Bundle to layer to finding to question, as the second pass will find it: 4 bundles · 8 layers · 8 findings · 16 questions.

Understand What a large language model is.

Science.

Definition

Definition.

Definition

Definition.

  1. What is a large language model, and how does it differ from earlier language models, search engines and general AI? definition
  2. Is the question about the technology in general, a specific product or a policy debate? boundary

Working

How it works.

Working

Working.

  1. How do tokens, transformers, pretraining, instruction tuning and feedback produce a usable model, in general terms? definition
  2. Which entry fits the specific component? action
Kinds Kinds and models.

Sources.

Architectures

Architectures.

Architectures

Architectures.

  1. How do dense, mixture of experts, multimodal and action-oriented models differ? provenance
  2. Which references are standard? provenance

Models

Notable models.

Models

Models.

  1. What are the major model families from different developers, described neutrally? provenance
  2. Is the presentation neutral across developers? boundary
Use Use and limits.

Practice.

Capabilities

Capabilities.

Capabilities

Capabilities.

  1. What can large language models do, and how are they evaluated? provenance
  2. Which sources are cited? provenance

Limits

Limitations.

Limits

Limits.

  1. What are hallucination, bias, prompt injection and other limitations, and how are they mitigated? provenance
  2. Which entry fits AI safety? action
Context Governance and history.

Context.

Governance

Governance.

Governance

Governance.

  1. What debates surround regulation, copyright, energy use and labour impacts, with positions attributed? provenance
  2. Is the presentation neutral? boundary

History

History.

History

History.

  1. How did language models develop from n-grams and RNNs to transformers and today s models? provenance
  2. Which entry fits the history of natural language processing? action

What the second pass must settle

  • Should transformer and instruction tuning be separate primary entries?
  • How should research sources be linked?
  • The registry entry has merged aliases naming specific products and architectures; should they be split off?