Tag: text-analytics

  • The Language of Surrender: Phrase Improvement Measurements

    The Language of Surrender: Phrase Improvement Measurements

    The trigger vocabulary that finds surrender statements was grown and tested: 373 mined candidates, 45 tested against ten labeled sources spanning four tiers of label quality, 9 promoted on human gold. Three measured regularities came out: precision rises with phrase length (0.603 at two words to 0.771 at five, over 134,600 decided spans); first-person forms separate self-acknowledged limitations from reviewer criticism (0.955 versus 0.000 for the same phrase family); and mining scores do not predict trigger quality. Every number is a measurement with its label tier stated.

  • The Language of Surrender: Larger Model Measurements

    The Language of Surrender: Larger Model Measurements

    Three models one size class up — Qwen3-14B, Phi-4, and Mistral-Small-24B at Q4_K_M quantization — measured on the same 1,505-sentence limitation gold set, same prompt, same harness as the earlier six-model comparison. Qwen3-14B scores F1 0.8272 as a single model, above the previous two-model union (0.8143) and the published rule-system reference (0.800). VRAM at load: 10.6 GB of 16.3 GB. Every number is a full-split measurement.

  • The Language of Surrender: Model Selection Observations

    The Language of Surrender: Model Selection Observations

    Scientific papers state their own defeats in words: a sample that could not be enrolled, a computation that could not be afforded, data that could not be obtained. A phrase list finds those statements with precision 0.23. Judging each match in context — with rules, two trained classifiers, and a pair of 8-billion-parameter models on two 16GB GPUs — raises the measured operating point to a 0.945-precision auto-admit gate and a two-judge F1 of 0.814, beating the published rule-based reference on the same gold set. Every number is a full-split measurement.

  • The Search for EULA Eve
    · rreck · data

    The Search for EULA Eve

    Population genetics has Mitochondrial Eve: the ancestor every living human's mitochondrial DNA traces back to. Fine print has an equivalent question. Legal text is almost never written fresh — it descends. This is a technical orientation on measuring that relatedness: content hashes for identity, winnowing fingerprints for inherited wording, on-prem SLM embeddings for kinship of meaning, and a dated derivation graph for descent — across 6,738 documents and 125,295 clauses. The search for the common ancestor turns up something real: the warranty disclaimer, carried by nearly every license measured.

  • Considerations in data sovereignty
    · rreck · data

    Considerations in data sovereignty

    Data sovereignty means continuing, revocable control over data about yourself. A measurement of 5,791 current agreements, 861,316 policy snapshots across 22 years, and 19 before-and-after GDPR pairs shows the mobile-app EULA is engineered to make that control impossible: postgraduate prose, terms that can be rewritten silently, no version to cite, and no way to say no.

  • Readability Score Distributions Across Project Gutenberg: A 2026 Extension

    Readability Score Distributions Across Project Gutenberg: A 2026 Extension

    In 2007, Reck & Reck computed seven readability measures for 15,511 Project Gutenberg texts and characterised the distribution of each across the corpus. This work extends that study to the present corpus — 60,787 English-prose works, the seven measures, lexical and structural features, lexicon-based sentiment, and emotional arcs — with author metadata drawn from Wikidata and an interactive tool for exploring any author against the whole distribution.

  • Written to Be Agreed To, Not Read
    · rreck · data

    Written to Be Agreed To, Not Read

    In 2013 a draft paper set out to measure the mobile-app agreements nobody reads, then stalled with its statistics left as 'XX' placeholders. This finishes it. A bibliometric analysis of 92 real mobile-app EULAs finds a postgraduate reading level, ~11 minutes each, and a structural trap: three-quarters can be rewritten at any time and bind you by 'continued use,' yet only 3% carry a version number. The agreements are written to be agreed to, not read — with the fine print, verbatim, to prove it.