zoukankan      html  css  js  c++  java
  • 综合练习:英文词频统计

    # -*- coding:UTF-8 -*-
    # -*- author:deng -*-
    news = '''
    The only problem unconsciously assumed by all Chinese philosophers to be of
    any importance is:How shall we enjoy life, and who can best enjoy life? No
    perfectionism, no straining after the unattainable, no postulating of he
    unknowable; but taking poor, modal human nature as it is, how shall we
    organize our life so that we can woke peacefully, endure nobly and live happily?
    Who are we? That is first question. It is a question almost impossible to answer.
    But we all agree with the busy self occupied in our daily activities is not quite
    the real self. We are quite sure we have lost something in the mere pursuit of living.
    When we watch a person running about looking for something in a field, the wise man can
    set a puzzle for all the spectator to solve: what has that person lost? Some one thinks
    is a watch; another thinks it is a diamond brooch; and others will essay other guesses.
    After all the guesses have failed, the wise man who really doesn't know what the person
    is seeking after, tells the company:" I'll tell you. He has lost some breath." And no one
    can deny that he is right. So we often forget our true self in the pursuit of living, like
    a bird forgetting its own danger in pursuit of a mantis which again forgets its own danger
    in pursuit of another.
    '''
    # 将分隔符替换为空格
    symbol=[",",".","!","?",":","'"]
    for i in range(len(symbol)):
    news=news.replace(symbol[i]," ")
    print(news)
    wordList = news.lower().split()

    # 将所有大写转换为小写
    news = news.lower()
    print(news)

    # 生成单词列表
    news = news.split()
    print(news)


    # 生成词频统计

    dict = {}
    for w in wordList:
    dict[w] = dict.get(w,0)+1
    print(dict)

    # 排除语法型词汇,代词、冠词、连词
    word = ['the', 'by', 'to', 'be', 'of', 'and', 'with', 'not']
    for i in word:
    del dict[i]
    print(dict)

    # 输出词频top10
    dict1 = sorted(dict.items(), key=lambda x: x[1], reverse=True)
    for i in range(10):
    print(dict1[i])
  • 相关阅读:
    open_basedir restriction in effect,解决php引入文件权限问题 解决方法
    Linux增加虚拟内存方法
    centos下kill、killall、pkill命令区别
    正则表达式全集
    Mysql中外键的 Cascade ,NO ACTION ,Restrict ,SET NULL
    配置frp实现内网穿透
    解决微信授权回调页面域名只能设置一个的问题 [php]
    高性能Mysql主从架构的复制原理及配置详解
    【Git】工作中99%能用到的git命令
    SVN服务器搭建和使用
  • 原文地址:https://www.cnblogs.com/dfq621/p/8655087.html
Copyright © 2011-2022 走看看