Urgent.News

What's breaking now, across thousands of outlets.

AI

Sample Efficiency: Humans vs Models

Everyone agrees models need far more data than children. The size of the gap depends entirely on decisions about what to count, and those decisions move the answer by more than the disagreement they are meant to settle. Why there is no single number The comparison sounds like it should reduce to a ratio: tokens seen by a model, words heard by a child, divide. Four things prevent that. The human…

Comparing a child's language development to that of a language model is often portrayed as a simple ratio of the amount of data each consumes. However, this comparison is far more complex than it appears. The ratio of tokens seen by a model to words heard by a child fails to capture the full picture due to several factors.

Firstly, the human figure used for comparison is an estimate from observational studies, which are subject to various inconsistencies. These studies vary not only between families but also in methodology, leading to significant discrepancies in the estimated figures. Moreover, a child's input is not merely words; it includes visual, auditory, and proprioceptive information, arriving in a causally structured stream that the child controls. Counting only the words misses out on the vast majority of the child's learning experience.

Secondly, the target for comparison is undefined. A five-year-old and a language model are not equally competent in the same ways, making any notion of "comparable competence" dependent on specific tasks. This lack of a clear benchmark further complicates the comparison.

Thirdly, the question of whether evolutionary optimization should be counted as training is a matter of debate. If this process is included, the human number becomes astronomically large, while the model appears more efficient. Conversely, if it is excluded, the model may look more efficient, but this perspective is controversial.

These factors illustrate that the sample efficiency gap between humans and models is not a single, quantifiable quantity. Instead, it depends on the specific counting choices made. The page quoted acknowledges these counting choices and provides five different ways to count, each yielding a different result.

To resolve the confusion, it is essential to recognize that these are five different quantities, not estimates of a single metric. Citing one of these figures without specifying the counting method contributes to the prevailing confusion. The BabyLM Challenge offers a promising approach to narrowing the gap by restricting training to a corpus on the scale of what a child might hear in the first years of life, then evaluating the resulting models on standard linguistic benchmarks.

This approach converts the rhetorical comparison into a shared, reproducible benchmark with a fixed budget, allowing researchers to explore how different techniques can help close the gap.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Schema Evolution in AI Pipelines

Upstream will add a field, rename a field, change a type from string to object, and start sending null where it never did. None of that is avoidable.

More from Wednesday 12 August →