Behavior Research Methods (2024)
The article Green, C., Keogh, K., Sun, H., & O’Brien, B. (2024). The Children’s Picture Books Lexicon (CPB-LEX): A large-scale lexical database from children’s picture books. Behavior Research Methods, 56(5), 4504-4521, available here (https://doi.org/10.3758/s13428-023-02198-y) presents CPB-LEX, a large-scale database of lexical statistics derived from children’s picture books age range 0–8 years.
CPB-LEX was built through an innovative method of computationally extracting lexical information from automatic speech-to-text captions and subtitle tracks generated from social media channels dedicated to reading picture books aloud. It consists of approximately 25,585 types, wordforms and their frequency norms, raw and Zipf-transformed, a lexicon of bigrams, and a document-term matrix.
Expanded Genre & Frequencies Update
This section provides an extended dataset to explore a large-scale lexical resource derived from children’s picturebook read-aloud videos publicly available online. It supplements the foundational database by substantially increasing corpus scale and introducing explicit genre modeling traits.
This data segment captures words and structural counts computed across approximately 5,874 read-aloud transcripts, expanding performance bounds to approximately 3.49 million word tokens and 44,617 word types. Additional structural matrices sorted across narrative and informational tracks are partitioned via Large Language Model classifications.
Green, C. & Keogh, K. (2026). A Dataset for Studying the Vocabulary of Picturebooks: A Supplement to the Children’s Picturebook Lexicon (CPB-Lex). Data in Brief.
Access the primary experimental data repository containing 25,585 types, Zipf norms, bigram metrics, and foundational document-term arrays.
Access the expanded 3.49 million token matrix, encompassing 5,874 read-aloud transcripts with LLM-generated narrative and informational genre vectors.
Data also hosted at Mendeley Data, DOI: 10.17632/9pgmnfkf56.1