Data Methodology
1. Job sources
Jobs come exclusively from public employer career pages served through public ATS APIs — currently Greenhouse and Lever. The employer's own posting is treated as the source of truth. We do not scrape gated sources or third-party aggregators.
2. Refresh frequency
Company boards are re-fetched on a rolling cycle (each company at least every ~12 hours when the sync worker runs continuously). Every job stores first_seen_at, last_seen_at and last_verified_at, shown on each job page as "Verified X hours ago".
3. Role classification
Job titles are lowercased, normalized and matched against ordered pattern rules per canonical role (30 roles, each with aliases). Matching stores the classification in the database; unmatched titles remain visible on the board but are excluded from role-level statistics. Role clusters group related titles (e.g. Forward Deployed Engineer / FDE / Deployment Engineer) so adjacent roles can be compared without pretending they're identical.
4. Skill extraction
Skills are detected by conservative pattern matching (word-boundary regexes) against job title and description text. Each match stores a confidence value and an evidence snippet. Ambiguous tokens are deliberately excluded (e.g. "rag" only matches as a standalone word, never inside "storage").
5. Salary methodology
Salary figures come only from ranges printed in the posting text, parsed with strict sanity bounds (plausible magnitude, min < max, ratio < 3×) and stored with source type job_description. We never mix in self-reported data, and we never publish medians below 5 observations.
6. Deduplication
Jobs are keyed by (source, external_id). Content hashes detect title/description changes between syncs. The same title at two locations is not merged — they are separate postings.
7. Job status lifecycle
- active — seen and verified in the most recent successful sync of its company
- stale — missed for 3+ days of successful syncs (grace period for transient failures)
- closed — unseen for 21+ days across repeated checks, or explicitly marked closed
Crawler errors never delete or close jobs — a failed sync leaves the last known state untouched and marks the company's sync as errored.
8. Limitations
The dataset covers tracked employers only — a lower bound on the market, skewed toward companies that use public ATS systems and toward the US market (salary disclosure follows US pay-transparency laws). Role and skill classification is rule-based and conservative; precision is prioritized over recall.
9. Corrections
Anyone can report an error from any page (see About). Corrections update the underlying dataset, not just the page.