The Children's Picturebook Lexicon (CPB-Lex)

An open-access empirical interface exploring vocabulary diversity, token frequencies, and genre distributions across childhood print exposure models.
About the Datasets & Publications (Click to expand)

Original CPB-LEX Publication

Behavior Research Methods (2024)

The article Green, C., Keogh, K., Sun, H., & O’Brien, B. (2024). The Children’s Picture Books Lexicon (CPB-LEX): A large-scale lexical database from children’s picture books. Behavior Research Methods, 56(5), 4504-4521, available here (https://doi.org/10.3758/s13428-023-02198-y) presents CPB-LEX, a large-scale database of lexical statistics derived from children’s picture books age range 0–8 years.

CPB-LEX was built through an innovative method of computationally extracting lexical information from automatic speech-to-text captions and subtitle tracks generated from social media channels dedicated to reading picture books aloud. It consists of approximately 25,585 types, wordforms and their frequency norms, raw and Zipf-transformed, a lexicon of bigrams, and a document-term matrix.

Data-in-Brief CPB-LEX Supplementary Data

Expanded Genre & Frequencies Update

This section provides an extended dataset to explore a large-scale lexical resource derived from children’s picturebook read-aloud videos publicly available online. It supplements the foundational database by substantially increasing corpus scale and introducing explicit genre modeling traits.

This data segment captures words and structural counts computed across approximately 5,874 read-aloud transcripts, expanding performance bounds to approximately 3.49 million word tokens and 44,617 word types. Additional structural matrices sorted across narrative and informational tracks are partitioned via Large Language Model classifications.

Green, C. & Keogh, K. (2026). A Dataset for Studying the Vocabulary of Picturebooks: A Supplement to the Children’s Picturebook Lexicon (CPB-Lex). Data in Brief.

Download the FULL CPB LEX

Access the primary experimental data repository containing 25,585 types, Zipf norms, bigram metrics, and foundational document-term arrays.

Open Core Repository on OSF ↗

Download UPDATED SUPPLEMENT (More Data)

Access the expanded 3.49 million token matrix, encompassing 5,874 read-aloud transcripts with LLM-generated narrative and informational genre vectors.

Data also hosted at Mendeley Data, DOI: 10.17632/9pgmnfkf56.1

Open Supplement Repository on OSF ↗ Open Mendeley Data DOI ↗
DOI link will resolve once the Mendeley Data record is live.
Core Database Matrices
Supplementary Dataset (Data-in-Brief Extended Pool)
Loading data...
Matching Rows: 0