← Back to catalogue
Research draft

information retrieval

vr.tr.information-retrieval · INF.KNW

Let an agent explain information retrieval, relay models, indexing, evaluation and applications from computer and information science sources, describe the named approaches and flag misfiled aliases, and distinguish information retrieval from data retrieval in databases, information extraction, sequential access in storage and interoperability standards.

Thing Registry Information and virtual systems

Research draft, second pass

A second pass drafted this model: the structure a model of this thing needs, and what is known about it in the world. The line under this one says how the second half was obtained - researched against sources, or recalled without web access, in which case nothing here was read anywhere and every claim is a lead to verify. Unreviewed either way.

written by Claude from model knowledge without web access - no source was read, every claim is a lead to verify

Researched by: Claude

Purpose and description

Let an agent explain information retrieval, relay models, indexing, evaluation and applications from computer and information science sources, describe the named approaches and flag misfiled aliases, and distinguish information retrieval from data retrieval in databases, information extraction, sequential access in storage and interoperability standards.

The activity and field of finding material, usually documents or data, that satisfies an information need from within large collections, using techniques such as indexing with inverted indexes, full-text search, ranking models such as TF-IDF and BM25, and increasingly neural and vector search, evaluated with measures such as precision and recall in benchmark efforts such as the TREC conferences begun in 1992, and applied in web search, library catalogues, computer-aided retrieval systems and geospatial information retrieval; the registry aliases data extraction, sequential access and health information interoperability name related but distinct concepts.

What it is for: Finding relevant information in collections.

It can be explain models and indexing; relay evaluation; describe named approaches; flag misfiled aliases.

Distinguishing features

Relevance ranking

Unstructured data

Evaluation benchmarks

Rapidly evolving methods

What it looks like

Not a visible object; seen as search boxes and ranked result lists.

Physical character

TREC conferences began: 1992 year - NIST

BM25: probabilistic ranking function note - Okapi system

SMART system: 1960s note - Gerard Salton

How it is recognised

Finding relevant information in collections

Data extraction, full-text search, sequential access, health information interoperability, computer-aided retrieval, geospatial information retrieval

Database queries return exact matches; information extraction pulls structured facts; sequential access is a storage method; interoperability concerns data exchange

Related models

is a kind of - in registry terms

search

uses -

inverted index

is evaluated by -

precision and recall

is contrasted with -

information extraction

In practice

Families and kinds

Boolean retrieval

vector space and probabilistic models

full-text search engines

neural and dense retrieval

geospatial and multimedia retrieval

retrieval-augmented generation

Standards and regulation

Z39.50 and SRU protocols for library search

Data protection laws affecting search logs

Failure modes and hazards

Treating relevance as truth

Bias in ranking

Misfiled aliases

Also called

data extractionfull-text searchsequential accesshealth information interoperabilitycomputer-aided retrievalGeospatial Information retrievalpollingfaceted searchprivate information retrievalrandom accessdata retrievalbrowsingGeographic information retrievalkeyword searchinglegal information retrievalpearl growinglinear searchcross-language information retrievalweb surfingweb navigationcitation snowballingweb crawlingcaret browsing

Where this came from

wikidata · CC0 1.0

Drafted structure

Bundle to layer to finding to question, as the second pass will find it: 4 bundles · 8 layers · 8 findings · 16 questions.

Understand What information retrieval is.

Definition.

Definition

Definition.

Definition

Definition.

  1. What is information retrieval, and how does it differ from database queries, information extraction, sequential access and interoperability? definition
  2. Is the question about search engines, libraries, research or building a system? boundary

Named

Named approaches and aliases.

Named

Named.

  1. What are full-text search, computer-aided retrieval and geospatial IR, and why are data extraction, sequential access and health interoperability different? definition
  2. Which entry fits the specific concept? action
Models Models and indexing.

Science.

Indexing

Indexing.

Indexing

Indexing.

  1. How do inverted indexes enable fast search? provenance
  2. Which references are standard? provenance

Ranking

Ranking models.

Ranking

Ranking.

  1. How do TF-IDF, BM25 and neural models rank documents? provenance
  2. Which sources are cited? provenance
Evaluation Evaluation.

Application.

Metrics

Metrics.

Metrics

Metrics.

  1. What do precision, recall and nDCG measure? measurement
  2. Which entry fits precision and recall? action

TREC

Benchmarks.

TREC

TREC.

  1. How did TREC shape IR evaluation? provenance
  2. Which entry fits Text Retrieval Conference? action
Context Applications and society.

Context.

Web search

Web search.

Web search

Web search.

  1. How do web search engines apply IR at scale? provenance
  2. Which entry fits search engine? action

Bias

Ranking bias.

Bias

Bias.

  1. What research exists on bias and fairness in search rankings, with findings attributed? provenance
  2. Is the presentation neutral and attributed? boundary

What the second pass must settle

  • The registry aliases sequential access and health information interoperability should be moved
  • How should IR research venues be linked?
  • Should neural retrieval be a separate entry?