Research towards Indian handwritten document analysis achieved increasing attention in recent years. In pattern recognition and especially in handwritten document recognition, standard databases play vital roles for evaluating performances of algorithms and comparing results obtained by different groups of researchers. For Indian languages, there is a lack of standard database of handwritten texts to evaluate performance of different document recognition approaches and for comparison purpose. In this paper, an unconstrained Kannada handwritten text database (KHTD) is introduced. The KHTD contains 204 handwritten documents of four different categories written by 51 native speakers of Kannada. Total number of text-lines and words in the dataset are 4298 and 26115, respectively. In most of text-pages of the KHTD contains either an overlapping or a touching text-lines and the average number of text-lines in each document on the database is 21. Two types of ground truths based on pixels information and content information are generated for the database. Providing these two types of ground truths for the KHTD, it can be utilized in many areas of document image processing such as sentence recognition/understanding, text-line segmentation, word segmentation, word recognition, and character segmentation. To provide a framework for other researches, recent text-line segmentation results on this dataset are also reported. The KHTD is available for research purposes.
Conference proceeding
A benchmark kannada handwritten document dataset and its segmentation
Proceedings, 11th International Conference on Document Analysis and Recognition, pp.141-145
International Conference on Document Analysis and Recognition (ICDAR 2011) (Beijing, China, 18/09/2011 - 21/09/2011)
2011
Metrics
36 Record Views
UN Sustainable Development Goals (SDGs)
This output has contributed to the advancement of the following goals:
Source: InCites
Abstract
Details
- Title
- A benchmark kannada handwritten document dataset and its segmentation
- Creators
- Ali Reza AlaeiP NagabhushanUmapada Pal
- Publication Details
- Proceedings, 11th International Conference on Document Analysis and Recognition, pp.141-145
- Conference
- International Conference on Document Analysis and Recognition (ICDAR 2011) (Beijing, China, 18/09/2011 - 21/09/2011)
- Publisher
- IEEE; Los Alamitos, California
- Number of pages
- 141-145
- Identifiers
- 2013; 991012821741302368
- Academic Unit
- Information Technology; Faculty of Science and Engineering; School of Business and Tourism; Faculty of Business, Law and Arts
- Language
- English
- Resource Type
- Conference proceeding