早口hayakuchi

Hayakuchi (早口, “fast talk”) · a corpus study of anime subtitles

Dialogue density in anime

Dialogue density is the number of spoken morae per minute of runtime. Across the Japanese subtitles of 1,635 anime franchises the median is 192 morae/min. Among the 100 most popular franchises it ranges from 136 (Vinland Saga) to 387 (The Disastrous Life of Saiki K.).

10.0 s

Ten seconds of runtime at each franchise's median density, one square per mora. The mora is the timing unit of Japanese: か, ん, っ and the long-vowel mark ー each count as one.

34,075 episodes 2,185 seasons 12.0 million lines 13,400 hours of runtime

Rankings

Franchises ranked by measure

Seasons are pooled into franchises. Popularity is the number of AniList users who list a title, taken from the franchise's most popular season. The default view is the 100 most popular franchises.

Most popular

median of selection

Highest

    Lowest

      Predictors

      Metadata predictors of dialogue density

      An ordinary least-squares regression of franchise density on AniList genres, source material, format and air year (n = 1,635) gives R² = 0.19. Genres alone give R² = 0.11 and source material 0.09; format and year add little.

      The largest coefficient is for light-novel source: +22.7 morae/min relative to manga (SE 2.3), with genre, format and year held fixed. Retention of first-person narration from the source text is one possible explanation; it is not tested here. Among genres, Comedy has a coefficient of +14.9 (SE 1.8) and Horror −21.6 (SE 4.2).

      Density is weakly correlated with popularity (Spearman ρ = 0.13) and uncorrelated with mean AniList score (ρ = 0.02).

      Adjusted coefficient, with 95% confidence interval Unadjusted difference in means

      Source material (reference: manga)

      Genre (reference: franchises without the genre)

      Table: coefficients

      AniList tags

      Tags are user-voted and more specific than genres. For each tag on at least 25 franchises (relevance of 60% or more), the chart shows the mean regression residual of the tagged franchises, that is, the association with density beyond genre, source, format and year, together with the unadjusted difference in medians.

      The largest positive residuals are for Meta (+18.9), Parody (+16.1) and Satire (+14.3); the largest negative for Iyashikei (−16.3), Steampunk (−13.2) and Fairy (−12.5). Tags overlap, so these estimates are not independent of one another.

      The 12 lowest and 12 highest of 142 tags by mean residual. Lines are 95% confidence intervals.
      Table: all tags

      Rate and coverage

      Speaking rate and dialogue coverage

      Dialogue density is the product of speaking rate, the morae per second while a dialogue caption is on screen, and dialogue coverage, the share of runtime with a dialogue caption on screen: density = 60 × rate × coverage. Caption timing conventions differ between subtitle sources, so this section uses only the 582 franchises with Netflix captions, which follow a single timing guideline.

      In a decomposition of the variance of log density across these franchises, coverage accounts for 55%, speaking rate for 27%, and their covariance for 20%. The Disastrous Life of Saiki K., the highest-density franchise in this subset, ranks 1st in speaking rate and 8th in coverage.

      Each point is a franchise. Curves are lines of equal density (morae/min). Highlighted: the 60 most popular franchises in the subset. Selecting a point opens it in the rankings.

      Within episodes

      Dialogue density by position within the episode

      Mean density in 15-second bins over 31,389 episodes with a runtime of 23 to 25 minutes. Lyrics are excluded from the counts, so opening and ending sequences appear as troughs. Density reaches its maximum of 229 morae/min at about minute 6 and declines gradually to about 205 by minute 20. Commercial breaks fall at different points in different series and leave no distinct trough in the mean. The small rise after the ending sequence is consistent with next-episode previews.

      Across seasons

      Variation within and between seasons

      Position within a season

      Each episode's density relative to the median of its season, over 1,548 seasons with complete episode numbering and at least 10 episodes. Medians for episodes 1 to 8 lie within 1% of zero. The penultimate episode is 2.4% below its season median and the final episode 3.8% below; 63% of finales fall below their season median.

      Table: by episode position

      Second seasons

      Across 341 franchises with at least two seasons, season 2 density correlates with season 1 density at r = 0.78, and the median change is +0.8%. Density is largely stable within a franchise.

      Production

      Studios and staff

      Median density by the main animation studio of each franchise's first season, for studios with at least 10 franchises. Lines show the interquartile range. Studio medians range from 164 (Toei Animation) to 226 (EMT Squared).

      Table: studios

      Directors and series-composition writers

      Median density for staff credited on at least 6 franchises, with each franchise counted once. Medians conceal wide individual ranges: Masaaki Yuasa has the lowest director median (158 morae/min), although his The Tatami Galaxy (296) is among the highest-density franchises in the corpus.

      Over time

      Change over time

      Seasons are dated by their own air year. The median of yearly medians is 185 morae/min for 2000–2009 and 196 for 2020 onward. Over the same period the share of seasons adapted from light novels, the source with the highest adjusted density, rose from 11% to 26%.

      In the franchise-level regression above, with genre and source held fixed, the year coefficient is +1.6 morae/min per decade (95% CI −0.2 to +3.4). The raw increase is therefore consistent with a change in the composition of what is produced rather than a general change in scripting.

      Density by air year

      Line: median season. Band: interquartile range. Years with at least 10 seasons.

      Share of seasons adapted from light novels

      Table: by year

      Vocabulary

      Vocabulary and register

      Vocabulary load is the number of the corpus's most frequent lemmas needed to cover 95% of a franchise's word tokens, excluding proper nouns and numerals. Lemma frequencies are averaged over franchises with equal weight. The median load is about 6,000 lemmas, and load is nearly uncorrelated with dialogue density (Spearman ρ = 0.09).

      The Tatami Galaxy and Food Wars! are high on both measures. The Disastrous Life of Saiki K. has the highest density among popular franchises but a below-median load. Girls' Last Tour is low on both.

      Highlighted: the 100 most popular franchises. Logarithmic vertical axis.

      First-person pronouns by demographic

      The choice of first-person pronoun in Japanese indexes gender and formality: 俺 (ore) is informal and masculine, 僕 (boku) masculine but softer, 私 (watashi) neutral or feminine, and あたし (atashi) informal and feminine. The chart gives each pronoun's share of first-person pronoun tokens, pooled over franchises with each AniList demographic tag. 俺 is the most frequent form in shounen titles (44%), and 私 in shoujo (48%), seinen (45%) and kids' titles (41%).

      Table: pronoun shares

      Polite forms and interjections

      The share of lines containing the polite auxiliaries です or ます is weakly and positively correlated with density (ρ = 0.28); polite forms are also longer in morae than their plain equivalents. The share of lines consisting only of interjections, such as えっ or うわっ, is weakly and negatively correlated (ρ = −0.22).

      Methods

      Data and methods

      • Corpus. Japanese subtitle files from the kitsunekko mirror, mostly streaming captions and Blu-ray transcripts. Japanese subtitles are used because English subtitles are condensed to meet reading-speed limits.
      • Sample. All series (TV, TV short or web) with an AniList ID and at least six usable episodes, one release per season and up to 26 episodes per season. Seasons are grouped into franchises through AniList prequel and sequel relations, except where a sequel introduces a new cast.
      • Cleaning. Song lyrics, on-screen signs, other-language lines, speaker labels, sound-effect captions and furigana are removed.
      • Measurement. Morae are counted from UniDic readings using the fugashi tokenizer. Dialogue density is morae divided by episode length from AniList, and a franchise's value is the median over its episodes.
      • Exclusions. AI-transcribed (Whisper) subtitles, which undercount by roughly 15 to 20% where they can be compared with human transcripts, and shorts under five minutes, for which AniList's whole-minute lengths are too coarse.
      • Subsets. Speaking rate and coverage use Netflix captions only. Vocabulary measures use UniDic lemmas and exclude proper nouns and numerals. Genres, tags and studios are taken from each franchise's first season.
      • Limitations. Subtitles approximate the audio, and overlapping or background speech may be missing. Some proper-noun readings are wrong, with negligible effect on counts. All results are observational associations.