A benchmark kannada handwritten document dataset and its segmentation

Ali Reza Alaei; P Nagabhushan; Umapada Pal

doi:10.1109/ICDAR.2011.37

Back

Conference proceeding

A benchmark kannada handwritten document dataset and its segmentation

Ali Reza Alaei, P Nagabhushan and Umapada Pal

Proceedings, 11th International Conference on Document Analysis and Recognition, pp.141-145

International Conference on Document Analysis and Recognition (ICDAR 2011) (Beijing, China, 18/09/2011 - 21/09/2011)

2011

DOI: https://doi.org/10.1109/ICDAR.2011.37

Files and links (1)

url

A benchmark kannada handwritten document dataset and its segmentationView

Published (Version of record)

Metrics

36 Record Views

23 Times Cited - Web of Science

UN Sustainable Development Goals (SDGs)

This output has contributed to the advancement of the following goals:

Source: InCites

Abstract

Business

Tourism and Travel

Research towards Indian handwritten document analysis achieved increasing attention in recent years. In pattern recognition and especially in handwritten document recognition, standard databases play vital roles for evaluating performances of algorithms and comparing results obtained by different groups of researchers. For Indian languages, there is a lack of standard database of handwritten texts to evaluate performance of different document recognition approaches and for comparison purpose. In this paper, an unconstrained Kannada handwritten text database (KHTD) is introduced. The KHTD contains 204 handwritten documents of four different categories written by 51 native speakers of Kannada. Total number of text-lines and words in the dataset are 4298 and 26115, respectively. In most of text-pages of the KHTD contains either an overlapping or a touching text-lines and the average number of text-lines in each document on the database is 21. Two types of ground truths based on pixels information and content information are generated for the database. Providing these two types of ground truths for the KHTD, it can be utilized in many areas of document image processing such as sentence recognition/understanding, text-line segmentation, word segmentation, word recognition, and character segmentation. To provide a framework for other researches, recent text-line segmentation results on this dataset are also reported. The KHTD is available for research purposes.

Details

Title: A benchmark kannada handwritten document dataset and its segmentation
Creators: Ali Reza Alaei
P Nagabhushan
Umapada Pal
Publication Details: Proceedings, 11th International Conference on Document Analysis and Recognition, pp.141-145
Conference: International Conference on Document Analysis and Recognition (ICDAR 2011) (Beijing, China, 18/09/2011 - 21/09/2011)
Publisher: IEEE; Los Alamitos, California
Number of pages: 141-145
Identifiers: 2013; 991012821741302368
Academic Unit: Information Technology; Faculty of Science and Engineering; School of Business and Tourism; Faculty of Business, Law and Arts
Language: English
Resource Type: Conference proceeding

A benchmark kannada handwritten document dataset and its segmentation

Files and links (1)

Metrics

UN Sustainable Development Goals (SDGs)

Abstract

Details

Southern Cross University Social media