Bijankhan Corpus
id:
bijankhan-corpus-192-11935726
title:
Bijankhan Corpus
text:
The Bijankhan corpus is a tagged corpus that is suitable for natural language processing (NLP) research on the Persian language. This collection is gathered from daily news and common texts. In this collection all documents are categorized into different subjects such as political, cultural, etc.; in about 4300 different subject categories. The corpus contains about 2.6 million manually tagged words with a tag set that contains 550 Persian part-of-speech tags. The Bijankhan corpus was created by
brand slug:
wiki
category slug:
encyclopedia
description:
original url:
https://en.wikipedia.org/wiki/Bijankhan_Corpus
date created:
date modified:
2023-10-10T22:52:37Z
main entity:
{"identifier":"Q4907174","url":"https://www.wikidata.org/entity/Q4907174"}
image:
{"content_url":"https://upload.wikimedia.org/wikipedia/commons/2/22/Bijankhan_Corpus_Logo.gif","width":117,"height":160}
fields total:
13
integrity:
14