KwickClips Artificial Intelligence · 46 sec · free
What comes after sentence segmentation?
Short answer. Tokenisation.
What is the last step?
Reducing words to roots.
Normalisation steps
| Sentences, then tokens |
| Remove stop words, symbols |
| Use lower case |
| Reduce words to roots |
Remember
| Segment, tokenise, remove |
| Common case, then root |
Before AI reads text, it cleans it. How? Four steps. Split text into sentences, then tokens. Remove stop words and symbols. Use lower case. Reduce words to roots. In order. Segment, tokenise, remove. Common case, then root.
This clip is from the full lesson: Text Processing: Normalisation, Tokenisation, Bag of Words and TF-IDF — 8 minutes, with the tables, the quick answers and the whole lesson in text.
Useful for: CBSE Class 10 Artificial Intelligence (417), CBSE Class 10 Artificial Intelligence (417), CBSE Class 10 Artificial Intelligence (417)
More KwickClips from this lesson
Voice-over is AI-generated; the script is written and checked by Kajal Ma'am. Confirm anything you plan around against your official board document. We never ask for a password or an OTP.




