- Data sources
- A search engine’s web page corpus
- Common Crawl Corpus
- The extraction method depends on where the number appears in the page
- Meta tags and microdata: structured data extraction
- Page header and footer: a rule-based extraction system
- Page body:
- Hand-engineered features for phone numbers (HTML tag structure, surrounding text, and so on)
- Machine learning for extraction and classification
1 min read