big data
Enable an agent to recognise a big-data collection in its processing context, judge whether it is fit for an intended use, and select feasible, permitted operations.
Research draft, second pass
A second pass drafted this model: the structure a model of this thing needs, and what is known about it in the world. The line under this one says how the second half was obtained - researched against sources, or recalled without web access, in which case nothing here was read anywhere and every claim is a lead to verify. Unreviewed either way.
recalled by Codex without web access - no source was read
Researched by: Codex
Purpose and description
Enable an agent to recognise a big-data collection in its processing context, judge whether it is fit for an intended use, and select feasible, permitted operations.
Big data comprises datasets whose volume, velocity, variety or complexity makes their effective storage, management and analysis require scalable approaches beyond conventional methods in the relevant operating context.
It can be Profile representative partitions and time windows to estimate scale, skew, missingness and schema variation.; Assess whether a proposed workload is feasible under stated latency, cost and resource limits.; Propose partitioning, incremental processing, sampling or aggregation with explicit consequences for fidelity.; Trace outputs to source versions, ingestion intervals and transformation steps.; Gate joins, exports and reuse against the applicable permissions and sensitivity of the combined data.; Plan retention, correction and deletion across replicas and derived datasets, then record verification evidence..
Distinguishing features
Identify a workload and resource baseline for which size, arrival rate or heterogeneity creates a demonstrated processing constraint; a large byte count alone does not establish the classification.
Distinguish the information collection from the data lake, database or cluster that stores or processes it; changing platforms need not create a different collection.
Distinguish streaming from big data: a low-rate stream may be readily handled, while a static collection may create substantial scale constraints.
Distinguish big data from AI training data: neither model training nor the presence of labels is required, and training datasets need not qualify as big data.
Treat usefulness, accuracy and representativeness as properties to investigate rather than consequences of having many records.
Scope
+ Collection boundaries, constituent datasets, streams and changing membership
+ Workload-relative evidence that the collection requires scale-aware processing
+ Formats, schemas, semantic variation and cross-source integration
+ Provenance, quality, representativeness and fitness for a stated use
+ Processing feasibility, access constraints and lifecycle obligations
- Data engineering and data science as disciplines or occupations
- Storage clusters, cloud services and processing engines as independently managed systems
- Analytical models, reports and decisions produced from the collection
- Individual source documents or creative works and their edition histories
- General organisational governance and legal regimes beyond constraints applicable to this collection
Characteristics
- Collection membership rule
- Snapshot, rolling window, append-only history or explicitly versioned source selection Defines which records an agent is assessing and whether repeated operations address the same collection.
- Logical and physical volume
- Records or events and bytes, with timestamp, compression and replication basis Separates information growth from storage overhead and supports capacity estimates.
- Arrival and mutation rates
- Events per second, bytes per second and updates per second over stated intervals Reveals whether ingestion and processing can keep pace with ordinary and peak demand.
- Structural and semantic heterogeneity
- Inventory of formats, schema versions, encodings, units and identifier systems Makes integration difficulty visible beyond raw size.
- Workload-to-resource relationship
- Collection linked to workload, resource baseline, service target and observed bottleneck Grounds the big-data designation in an explicit operational context.
- Freshness and completeness
- Event-to-availability delay and received-to-expected coverage, where an expected population is known Distinguishes current-looking output from output based on complete, timely input.
- Fitness for intended use
- Unassessed, supported, conditionally supported or unsupported for a named use Prevents quality judgments for one workload from being applied automatically to another.
- Permitted-use envelope
- Collection segments linked to evidenced access, purpose, location, sharing and retention constraints Determines which operations an agent may perform across sources with different restrictions.
Also called
Where this came from
wikidata · CC0 1.0
Drafted structure
Bundle to layer to finding to question, as the second pass will find it: 6 bundles · 11 layers · 17 findings · 27 questions.
Collection identity and scale Establishes which information belongs to the collection and why its scale matters in a specified context.
Big data has neither a useful universal size threshold nor a stable identity unless membership and workload assumptions are explicit.
Membership and versions
Defines source coverage, record inclusion and the identity of changing collection states.
Bounded collection
Record inclusion rules and reproducible references to snapshots or stream intervals.
- Which sources, record types, time windows and exclusions define this collection? boundary
- How can an agent identify the exact snapshot or stream interval used by an earlier operation? provenance
Scale in context
Connects measured data demands to processing limits.
Demonstrated scale constraint
Record the workload, baseline resources and measured constraint supporting the designation.
- Which workload exceeds a stated capacity, latency or cost target on the available baseline? measurement
- Is the limiting factor volume, arrival rate, heterogeneity, data skew or an evidenced combination? definition
Ingestion and temporal behaviour Describes how records arrive, change and become available for use.
At scale, backlog, duplicate delivery and late events can change the meaning of apparently current results.
Flow and backlog
Measures incoming demand against sustained processing capacity.
Ingestion capacity
Record typical and peak rates, processing throughput and outstanding work.
- What are the typical and peak arrival rates, and over which observation windows were they measured? measurement
- When input exceeds processing capacity, should the system buffer, throttle, sample or reject records, and what loss is permitted? action
Event time and revisions
Distinguishes event occurrence, ingestion, availability and subsequent correction.
Late, duplicate and corrected records
Record temporal semantics and rules for revising results as input changes.
- Which timestamps represent occurrence, ingestion and availability, and how are their uncertainties recorded? definition
- How should late arrivals, repeated deliveries, out-of-order events and source corrections update prior outputs? action
Heterogeneity and lineage Makes diverse source representations interpretable and transformations traceable.
Combining large sources requires evidence that fields, entities and versions remain compatible throughout processing.
Schema and semantic alignment
Records differences that prevent direct combination of sources.
Cross-source compatibility
Identify schema drift, conflicting meanings and uncertain entity matches before integration.
- Which schema versions, units, encodings or field meanings differ across sources and time periods? definition
- Which identifiers support joins, and how are collisions, missing keys and uncertain matches handled? boundary
Source-to-output traceability
Connects collection segments and outputs to their origins and transformations.
Lineage at usable resolution
Record enough lineage to investigate errors and reproduce outputs without assuming record-level tracing is always available.
- Can an output be traced to source versions, partitions, ingestion intervals and transformation configurations? provenance
- What lineage resolution is needed to isolate a faulty source and reprocess only the affected data? action
Quality and inference limits Assesses defects, coverage and the limits of conclusions drawn from the collection.
Large record counts can conceal systematic omissions, correlated observations and concentrated errors.
Partition-aware quality
Examines quality across sources, groups and time rather than only through collection-wide averages.
Defect distribution
Record where missingness, duplication, invalid values and drift occur and whether profiling can detect them.
- How do missingness, duplicates and invalid values vary by source, partition and time window? measurement
- Which checks require a full scan, and which sampling methods adequately cover rare or high-risk segments? action
Coverage and valid use
Relates observed records to the population or process an intended analysis concerns.
Representation and dependence
Record collection mechanisms, exclusions and repeated observations that limit inference.
- Which people, events or conditions are omitted or overrepresented by the collection mechanism? boundary
- How many records represent repeated or dependent observations rather than distinct units relevant to the intended analysis? measurement
Processing and controlled lifecycle Connects feasible processing choices with constraints on use and persistence.
Distributed processing and extensive copying affect both operational feasibility and the ability to enforce collection-specific obligations.
Execution and fidelity
Evaluates the consequences of processing plans for performance and result meaning.
Feasible processing plan
Record resource estimates and the fidelity implications of partitioning, approximation or incremental execution.
- What scan volume, data movement, partition skew, latency and cost are expected for the proposed workload? measurement
- Can sampling, aggregation or incremental processing meet the use requirement, and what errors or omissions must accompany the result? action
Use rights and propagation
Tracks permitted operations and lifecycle changes across combined sources and derived copies.
Permitted combination and retirement
Record evidence for use restrictions and mechanisms for propagating correction, expiry and deletion.
- What evidence permits the proposed access, linkage, processing location and export for every contributing source? provenance
- How will corrections, retention expiry and deletion requirements reach replicas, caches and derived datasets, and how will completion be verified? action
Evidence and external alignment What the world already says about this thing, gathered so the model can be checked against it.
A model that cannot be lined up against existing standards, identifiers and practice cannot be adopted by anyone who already uses them.
Reported evidence
Findings from the breadth pass, kept separate from the structural claims.
Check these first
Recalled without web access and unsourced; every item is a lead to verify.
- This describes the general information-technology sense; no narrower registry sense was supplied.
- The listed kinds overlap: structure and processing mode are separate classification dimensions.
- This is recall rather than source-checked research; standards editions and the registry's intended classification should be checked first.
- Which of these check these first hold for the sense of big data this model covers, and on what evidence? provenance
Kinds and varieties
Recalled without web access and unsourced; every item is a lead to verify.
- Structured big data
- Semistructured big data
- Unstructured big data
- Streaming big data
- Batch big data
- Which of these kinds and varieties hold for the sense of big data this model covers, and on what evidence? provenance
Standards and regulation
Recalled without web access and unsourced; every item is a lead to verify.
- ISO/IEC 20546, Information technology - Big data - Overview and vocabulary, issued by ISO and IEC.
- NIST Big Data Interoperability Framework, published by the US National Institute of Standards and Technology.
- Which of these standards and regulation hold for the sense of big data this model covers, and on what evidence? provenance
Real-world use
Recalled without web access and unsourced; every item is a lead to verify.
- Analyzing transaction streams to detect fraud and anomalies.
- Processing equipment telemetry for predictive maintenance.
- Analyzing large scientific datasets, including astronomical observations and genomic sequences.
- Building recommendation systems from large collections of interaction records.
- Analyzing service logs and network events for operational monitoring.
- Which of these real-world use hold for the sense of big data this model covers, and on what evidence? provenance
Typical measurements
Recalled without web access and unsourced; every item is a lead to verify.
- Data volume - No universal threshold; depends on available infrastructure and workload. - bytes and multiples such as TB or PB
- Ingestion rate - Workload dependent; no defining minimum. - events/s or bytes/s
- Processing latency - Application dependent, from real-time processing to long-running batch jobs. - seconds
- Which of these typical measurements hold for the sense of big data this model covers, and on what evidence? provenance
Failure modes and hazards
Recalled without web access and unsourced; every item is a lead to verify.
- Missing, duplicated, inconsistent or poorly documented records undermine results.
- Large sample sizes can conceal selection bias, confounding and spurious correlations.
- Linking datasets can expose sensitive information or enable reidentification.
- Distributed processing can lose, duplicate or reorder events unless failures and retries are handled correctly.
- Storage, computation and data movement can incur substantial financial and energy costs.
- Which of these failure modes and hazards hold for the sense of big data this model covers, and on what evidence? provenance
Neighbouring kinds and how to tell them apart
Recalled without web access and unsourced; every item is a lead to verify.
- Data analytics - Data analytics is the activity of extracting insights from data; it can operate on small datasets as well as big data.
- Database - A database is an organized collection of data; being a database does not establish the scale or processing demands associated with big data.
- Data lake - A data lake is a storage architecture for diverse data; big data describes characteristics and processing demands rather than a particular architecture.
- Machine learning - Machine learning develops models from data; it neither requires big data in every application nor encompasses all big-data processing.
- Work, medium or creative form - Big data is a category of data and associated processing challenges, rather than an authored work, edition or carrier; individual datasets may have those additional attributes.
- Which of these neighbouring kinds and how to tell them apart hold for the sense of big data this model covers, and on what evidence? provenance
What the second pass must settle
- Does the registry intend big data to denote information collections, the broader technological phenomenon or a field of practice, given that no definition is recorded?
- Which existing Vercy data-collection, dataset or world models already own parts of this scope and should be referenced rather than duplicated?
- Which authoritative definition should anchor the model, and how should workload-relative scale be represented without imposing an unsupported universal threshold?
- What evidence is sufficient to classify a collection as big data when its processing constraints have not yet been benchmarked?
- Which lineage resolution, quality tolerances and lifecycle verification requirements should vary by intended use rather than be mandatory for every instance?