ScourJet
How it works Providers Use cases Compare Pricing Docs GitHub
Sign in Start free
Open source, provider-agnostic, and deterministic by design.
Describe the dataset. Watch it get built.
Pick a template for structure, write a prompt for intent. ScourJet compiles both into a discovery plan you approve, runs it in visible budget-capped cycles, and learns from your feedback. No black-box agent — every request is logged, priced and replayable.
Start free → Self-host the community edition
git clone scourjet/scourjet · docker compose up
How it works
A pipeline you can read, not an agent you have to trust.
STEP 1
Template + prompt
Structure from a template, intent from your prompt.
STEP 2
Search & discover
Query variants fan out; sources are trust-weighted.
STEP 3
Extract
Schema-guided rows — never invented, always sourced.
STEP 4
Score
Dedupe, validate and grade every row 0–100.
STEP 5
Learn & iterate
Flag rows; the next cycle re-ranks from your feedback.
No run without approval
Every request is logged with its price. Replay any cycle.
Hard budget caps
Per project and per month — worst case is bounded.
Quality you can audit
Per-row scores, per-field validation, source provenance.
Providers
Bring your own keys.
Optimized for Ujeebu out of the box — search, scrape and extraction in one call. Connect Apify, Bright Data or any endpoint that speaks the open provider spec.
U
Ujeebu
optimized
A
Apify
B
Bright Data
S
SerpAPI
Z
Zyte
+
Custom
Compare
Different tools, different jobs.
Clay enriches lists you already have. Apify runs scrapers you build or buy. Browse AI watches pages for changes. ScourJet starts from a sentence and ends with a scored, exportable dataset — with you approving every step.
“Why not just ask an LLM for the data?”
An LLM invents
Ask for 300 dentists and you'll get plausible names and numbers — some real, some confidently made up. ScourJet only extracts from pages it actually fetched. The LLM plans the run; it never writes a row.
An LLM is frozen in time
Its training data ends months before you ask. ScourJet reads the live web at run time — today's prices, this week's events, the clinic that opened last month.
An LLM can't be audited
No source, no validation, no way to check field by field. Every ScourJet row links to its source URL, with validated phone, email and a 0–100 quality score.
ScourJet
Clay
Apify
Browse AI
Built for
Prompt → dataset from the open web
Enriching GTM lists
Running scrapers (actors)
Monitoring pages
Plan approval before any spend
Hard budget cap per run
≈ credit limits
≈ usage limits
≈ plan quotas
Per-row quality scoring
Feedback-driven cycles
≈ retraining
Self-hostable, open source
≈ open SDK
Pricing
PAYG credit packs
Subscription credits
Subscription + usage
Subscription tiers
ScourJet vs Claydeep dive → ScourJet vs Apifydeep dive → ScourJet vs Browse AIdeep dive →
positioning snapshot — verify details against each vendor's current docs
Pricing
One product. Two ways to run it.
No subscriptions — buy credit when you need it.
Community
$0self-hosted · Apache-2.0
Full pipeline — plan, run, score, learn
Your provider keys, your infrastructure
Local LLM support (Ollama-compatible)
Same dashboard, no billing screens
Clone on GitHub
hosted · pay as you go
Hosted
$0/mo · 500 free credits on signup
$10 · 1,000 cr $50 · 5,500 cr $200 · 24,000 cr
Credits never expire — bigger packs cost less per credit
Provider costs passed through at cost, no markup
Hosted runners, schedules and webhooks
No card required to start
Start free →
Start in minutes
From a sentence to a scored dataset.
Run it free and self-hosted with your own keys, or let us host it. Either way you approve every step and own every row.
Open the app → Read the docs
$ get started
git clone scourjet/scourjet
docker compose up
# ready on :8080 ✓
No card required · Apache-2.0
ScourJet — open web data discovery · powered by Ujeebu
Use cases Docs GitHub Provider spec Open app