论文标题

单词级别频率关系的两个参数方程

A Two Parameters Equation for Word Rank-Frequency Relation

论文作者

Ding, Chenchen

论文摘要

令$ f(\ cdot)$为单词的绝对频率,而$ r $为频率下降顺序的单词等级,那么以下函数可以适合级别频率关系\ [f(r; s,t)= \ left(\ frac {\ frac {r _ {\ tt max}}}} {r}} {r} {r} {r} {1-s} \ left(\ frac {r _ {\ tt max}+t \ cd \ cdot r _ {\ tt exp}}} {r+t \ cdot r _ {\ tt exp}}} \ right)分别对等级的期望; $ s> 0 $和$ t> 0 $是从数据估算的参数。在行为良好的数据上,应该有$ s <1 $和$ s \ cdot t <1 $。

Let $f (\cdot)$ be the absolute frequency of words and $r$ be the rank of words in decreasing order of frequency, then the following function can fit the rank-frequency relation \[ f (r;s,t) = \left(\frac{r_{\tt max}}{r}\right)^{1-s} \left(\frac{r_{\tt max}+t \cdot r_{\tt exp}}{r+t \cdot r_{\tt exp}}\right)^{1+(1+t)s} \] where $r_{\tt max}$ and $r_{\tt exp}$ are the maximum and the expectation of the rank, respectively; $s>0$ and $t>0$ are parameters estimated from data. On well-behaved data, there should be $s<1$ and $s \cdot t < 1$.

扫码加入交流群

加入微信交流群

微信交流群二维码

扫码加入学术交流群,获取更多资源