Tag: corpus-linguistics

  • The Language of Surrender: Phrase Improvement Measurements

    The Language of Surrender: Phrase Improvement Measurements

    The trigger vocabulary that finds surrender statements was grown and tested: 373 mined candidates, 45 tested against ten labeled sources spanning four tiers of label quality, 9 promoted on human gold. Three measured regularities came out: precision rises with phrase length (0.603 at two words to 0.771 at five, over 134,600 decided spans); first-person forms separate self-acknowledged limitations from reviewer criticism (0.955 versus 0.000 for the same phrase family); and mining scores do not predict trigger quality. Every number is a measurement with its label tier stated.

  • The Language of Surrender: Model Selection Observations

    The Language of Surrender: Model Selection Observations

    Scientific papers state their own defeats in words: a sample that could not be enrolled, a computation that could not be afforded, data that could not be obtained. A phrase list finds those statements with precision 0.23. Judging each match in context — with rules, two trained classifiers, and a pair of 8-billion-parameter models on two 16GB GPUs — raises the measured operating point to a 0.945-precision auto-admit gate and a two-judge F1 of 0.814, beating the published rule-based reference on the same gold set. Every number is a full-split measurement.

  • Readability Score Distributions Across Project Gutenberg: A 2026 Extension

    Readability Score Distributions Across Project Gutenberg: A 2026 Extension

    In 2007, Reck & Reck computed seven readability measures for 15,511 Project Gutenberg texts and characterised the distribution of each across the corpus. This work extends that study to the present corpus — 60,787 English-prose works, the seven measures, lexical and structural features, lexicon-based sentiment, and emotional arcs — with author metadata drawn from Wikidata and an interactive tool for exploring any author against the whole distribution.