Methodology
How we source, verify, and publish local information for Oneida and Herkimer Counties.
Data Pipeline
Every business listing and event follows a structured pipeline from official source to published record. Nothing is published without verification.
1. Source Selection
We use official government datasets and institutional feeds. Examples include the NYS Liquor Authority Active Licenses database, public library event calendars, and municipal records. All sources must be publicly accessible and authoritative.
2. Import & Normalization
Raw data is imported and normalized into our database schema. Each field stores its source URL, discovered timestamp, and confidence level. No data is modified without recording the provenance of that change.
3. Entity Resolution
Multiple sources may reference the same business. Our entity resolution system detects potential duplicates using name similarity, address matching, and geographic proximity. Matches are reviewed before merging.
4. Verification
Records are verified for completeness and accuracy. We check for required fields, validate geographic data, and confirm the source citation is accessible. Only verified records are marked for publication.
5. Publication
Verified records are published to the public site. Each listing includes source attribution and timestamps. Published records remain linked to their provenance data.
Provenance Tracking
Every field in our database stores its provenance. This includes the source URL or dataset ID, the timestamp when it was discovered, when it was last verified, and a confidence level based on source trustworthiness.
When data conflicts between sources, we apply a priority order based on authoritativeness. Official government sources take precedence over aggregators. Direct institutional feeds take precedence over third-party listings.
What We Don't Do
- No scraping: We don't scrape Google, Yelp, Facebook, or other platforms that prohibit data extraction
- No invention: We never generate or invent facts. Missing data remains missing
- No purchased lists: We don't buy business databases with unclear provenance
- No robots.txt violations: We respect crawl policies and Terms of Service
Database Architecture
Our PostgreSQL database uses a structured schema with 12 entity types. Core entities include businesses, events, organizations, jobs, and grants programs. Each entity type links to provenance records that store source attribution.
The system includes entity resolution tables for duplicate detection, a redirect map for legacy URL migration, and change tracking for observed updates like business openings, closures, or moves.
Technical details: Our codebase is open-source and available for review. The database schema, import pipelines, and entity resolution algorithms are documented in the repository.