- 1、本文档共54页,可阅读全部内容。
- 2、有哪些信誉好的足球投注网站(book118)网站文档一经付费(服务费),不意味着购买了该文档的版权,仅供个人/单位学习、研究之用,不得用于商业用途,未经授权,严禁复制、发行、汇编、翻译或者网络传播等,侵权必究。
- 3、本站所有内容均由合作方或网友上传,本站不对文档的完整性、权威性及其观点立场正确性做任何保证或承诺!文档内容仅供研究参考,付费前请自行鉴别。如您付费,意味着您自己接受本站规则且自行承担风险,本站不退款、不进行额外附加服务;查看《如何避免下载的几个坑》。如果您已付费下载过本站文档,您可以点击 这里二次下载。
- 4、如文档侵犯商业秘密、侵犯著作权、侵犯人身权等,请点击“版权申诉”(推荐),也可以打举报电话:400-050-0827(电话支持时间:9:00-18:30)。
查看更多
* Bakeoff 2007 – 法国电信北京研发中心 Problems of NER with only local information “Many empirical approaches…make decision only on local context for extract inference, which is based on the data independent assumption. But often this assumption does not hold because non-local dependencies are prevalent in natural language.” Observation from Experiments: There are many seen named entities are missed; At least 10% of unseen and missed named entities have been labeled out correctly for at least once. “If the context surrounding one occurrence of a token sequence is very indicative of it being an entity, then this should also influence the labeling of another occurrence of the same token sequence in a different context that is not indicative of entity”. * Bakeoff 2007 – 法国电信北京研发中心 * Bakeoff 2007 – 法国电信北京研发中心 Local Features Unigram:Cn(n=-2,-1,0,1,2) Bigram:CnCn+1(n=-2,-1,0,1) and C-1C1 0/1 Features Assign 1 to all the characters which are labeled as entity and 0 to all the characters which are labeled as NONE in training data. In such way, the class distribution can be alleviated greatly , taking Bakeoff 2006 MSRA NER training data for example, if we label the corpus with 10 classes, the class distribution is: 0.81(B-PER), 1.70(B-LOC), 0.95(BORG), 0.81(I-PER), 0.88(I-LOC), 2.87(I-ORG), 0.76(EPER), 1.42(E-LOC), 0.94(E-ORG), 88.86(NONE) if we change the label scheme to 2 labels(0/1), the class distribution is: 11.14 (entity), 88.86(NONE) * Bakeoff 2007 – 法国电信北京研发中心 Non-local Features Token-position features(NF1) These refer to the position information(start, middle and last) assigned to the token sequence which is matched with the entity list exactly. These features enable us to capture the dependencies between the identical candidate entities and their boundaries. Entity-majority features(NF2) These refer to the majority label assigned to the token sequence which is matched with the entity list exactly. These features enable us to capture the dependencies between the identical e
您可能关注的文档
- 中学生礼仪主题班会.ppt
- 中学生禁毒主题班会.ppt
- 中学生英语语法总结.ppt
- 中学生行为规范主题班会.com].ppt
- 中学生青春期恋爱主题班会.ppt
- 中学生饮食习惯调查.ppt
- 中学生饮食调查与建议结题报告.ppt
- 中学防艾培训讲稿.ppt
- 中小企业发展、上市融资问题分析及整合对策.ppt
- 中小企业如何盈利(潘皖江)-中华讲师网.ppt
- 陕西省2024九年级化学上册第六单元碳和碳的氧化物专训5常见气体的制取的一般思路和方法课件新版新人教版.pptx
- 居家健康知识培训课件.pptx
- 陕西省2024九年级化学上册第七单元燃料及其利用课题1燃料的燃烧第1课时燃烧的条件灭火的原理和方法课件新版新人教版.pptx
- 凤栖梧-柳永-精品.pptx
- 居家家具知识培训课件.pptx
- 陕西省2024九年级化学上册第七单元燃料及其利用实践活动6调查家用燃料的变迁与合理使用课件新版新人教版.pptx
- 陕西省2024九年级化学上册第七单元燃料及其利用整合练课件新版新人教版.pptx
- 陕西省2024九年级化学上册第七单元燃料及其利用实验活动4燃烧条件的探究课件新版新人教版.pptx
- 《客户服务管理》课件.ppt
- 第12课 水陆交通的变迁课件(共31张PPT)2024-2025学年高二历史统编版选择性必修二.pptx
文档评论(0)