Dataset
Thai G2P v4 dataset is a Thai grapheme-to-phoneme dataset that was built from Thai W2P and Wiktionary th-pron transliterator. We split the dataset by using the first 2 characters for the split group.
Huggingface dataset: https://huggingface.co/datasets/pythainlp/thai-g2p-v4-dataset
Author: Wannaphong Phatthiyaphaibun
sources
- Thai W2P: https://huggingface.co/datasets/wannaphong/thai-w2p
- Wiktionary th-pron transliterator: PyThaiNLP/pythainlp#1437