AsoSoft text corpus

id: asosoft-text-corpus-278-12868178
title: AsoSoft text corpus
text: The AsoSoft text corpus is the first large-scale Kurdish text corpus, collected and processed by the AsoSoft research and development group. It contains 458,000 documents that are collected from sources such as websites, news agencies, books, and magazines. The corpus is partially tagged by topic, so it can be used for topic identification tasks. Also, it is applicable for extracting language model and computational lexicon information. Part of the corpus is available online for non-commercial u
brand slug: wiki
category slug: encyclopedia
description: Kurdish text corpus
original url: https://en.wikipedia.org/wiki/AsoSoft_text_corpus
date created:
date modified: 2023-11-24T18:09:45Z
main entity: {"identifier":"Q31794083","url":"https://www.wikidata.org/entity/Q31794083"}
image:
fields total: 13
integrity: 14

Related Entries

Explore Next Part