Tokenize Words
Checking your account…
Sign in to save your code and progress across devices. The lesson and problem statement remain public.
Loading the interactive Practice workspace.If it does not appear, the problem and learning material remain readable, but browser execution is unavailable.Reload Practice workspace
Problem
Implement tokenize(text). Return whitespace-separated lowercase word tokens.
Starter code
def tokenize(text):
passTest cases
sample
{
"args": [
"One TWO\nthree"
]
}Expected: ["one","two","three"]
Wizard outline
- Step 1: Normalize token case
Produce a lowercase view of the input before token boundaries are considered. Case normalization and token separation are independent transformations and are easier to reason about separately.
- Step 2: Split every whitespace boundary
Return lowercase tokens with empty whitespace fields removed. split() without an argument recognizes all whitespace and naturally returns an empty list for whitespace-only input.
Footguns and prerequisites
- split(' ') creates empty tokens and ignores tabs or newlines.
- strings
Reviewed references
Recommended approach and implementation
Lowercase once and use whitespace-aware split.
Why it works: The default split recognizes all whitespace runs and emits no empty tokens.
def _lower_text(text):
return text.lower()
def tokenize(text):
return _lower_text(text).split()