jobdesk Datafarm Spiders

Sign in

Enter your email address to receive a one-time sign-in code.

Check your email

We sent a 6-digit code to

jobdesk Datafarm Spiders
Dashboard Companies Live Task spiders
Users
Activity Feedback Admin Plugin
v1.0.0
Live– Jobs– Analysing– Failing– Needs config– Crawled today– New/Completed– 🔔 Alerts– ⧉ Dup URLs– 💾 Spool–
🕷
Loading Datafarm Spiders…

Add company

Paste a careers/jobs URL — or a bare company domain and let the system find it. It auto-detects the ATS, builds and verifies the crawler config against the live page, and starts. Everything below is optional.

Name and source key are auto-derived from each host.
Applies to the whole batch in Bulk mode
Advanced options
Unique id · auto-derived if blank
1440=daily · 60=hourly · 5=board
Claude + Gemini ceiling
Job data output
Fetch escalation — only if auto-detect fails

Edit company

These 3 override the "Proxy Tier" field in Edit Config for their stage when checked — the stored config is left as-is, this just wins at fetch time.

Edit crawl configuration

📘 How configs work

Edit the JSON configuration directly. The object must be valid JSON and will be validated on save.


          
        

📝 Report an Issue

Describe the problem so the self-healing pipeline can attempt a fix. Common example: "Spider stops paginating after page 1 — no errors thrown."

Paste the <a> tag or its parent element. The healing pipeline uses this as the strongest possible signal to derive the correct jobLinkSelector.

📸 A fresh HTML snapshot of the career page will be captured automatically when you submit. This gives the healing pipeline the current DOM structure to generate accurate selectors.

⚙ Default Crawl Configuration

This configuration is seeded into every new company spider at creation time. It is not applied to existing companies. Edit any default values you want new spiders to start with.

Copy Spider

Set Status

Manually override the company status. Use with care — workers will act on the new status immediately.

Review & edit analysis prompt

Edit the user message that will be sent to Claude. Click Save & Approve to save your edits and send, or Approve as-is to send the original without changes.

Pick Selectors

Click a mode button, then click on the page below.

Add Bulk

Saved sheets
Loading…

Paste rows from a spreadsheet (tab-separated columns, one row per line) into any cell, or add/edit rows by hand — or paste/import a whole sheet at once below. Use each column header's dropdown to say what it is. Click a row's ▶ to create just that one spider, or Import all. Save a sheet by name to come back to it later — open several at once as tabs.

Row defaults

Applied to every row that doesn't set its own value for that column.

A small AI call per row catches a company that already exists under a different domain, and tags a known corporate group (e.g. "Smile" → "Baloise"). Uncheck for a sheet you already know is clean, to skip the cost.

Schedule Claude Code Import

A Claude Code routine will create spiders from this sheet's rows at the given time.

Harvest by Keyword

Describe an industry or company set in your own words — e.g. "construction industry Switzerland, top companies by size" or "Swiss hospitals and clinics". A background service searches via SERP, extracts real company names, and resolves each one's careers page — same pipeline as the "Discover" button, just seeded from a description instead of one name.

About this table

Each row is one spider — one company's careers page and its crawl configuration.

Name
The company's display name.
Source Key
The company's canonical domain — the identity its job data is published under. One spider per source key.
Career URL
The page the spider crawls from (opens in a new tab).
ATS
The applicant-tracking-system platform detected on this site, if any (e.g. Greenhouse, Workday) — informational.
Interval
How often the listing is re-checked for new/removed jobs.
Last Crawled
When the most recent crawl attempt ran, successful or not.
Fails
Consecutive failed crawls in a row — resets to 0 on the next success.

Status

New
Just created — queued for analysis, no working config yet.
Active
Has a working, verified config — crawling on its normal schedule.
Stale
Hasn't produced a successful crawl in a while; still retrying on schedule.
Failing
Recent crawls are erroring or returning zero jobs — eligible for autonomous self-healing.
Suspended
Manually paused — excluded from the crawl schedule until reactivated.
Needs Config
Automatic analysis couldn't produce (or verify) a working config — needs a human fix via the dashboard's Pick Selectors or the recorder plugin.

See the full spider & configuration guide for how spiders work end to end.

Confirm