What You'll Build
An automated job search pipeline that runs twice a day on your Azure VM. It scrapes LinkedIn and Adzuna for job listings matching your keywords, filters them with AI, tracks results in a Google Sheet, and sends a Telegram digest.
The flow:
- Linux cron fires at the times you choose (e.g. 11 AM and 3 PM)
- A Python scraper calls Apify (LinkedIn) and Adzuna (job boards) for all your keywords
- The scraper pre-filters results: removes duplicates, checks title relevance
- The scraper saves everything to a JSON file on disk
- A bash wrapper triggers your OpenClaw agent
- The agent reads the JSON, does AI-powered filtering, deduplicates against your Google Sheet, writes new jobs, and sends a Telegram digest
Why this architecture? The scraper runs outside OpenClaw because API calls take 10-20 minutes. Running them inside the agent would burn tokens while waiting. The agent only sees pre-filtered results, so it processes them in seconds instead of minutes.
Estimated monthly cost:
| Component | Monthly |
|---|---|
| Apify LinkedIn scraper (Starter plan) | $29 |
| Adzuna API | Free |
| AI filtering tokens (Claude or GPT) | ~$15-30 |
| Google Sheets API + Telegram | Free |
| Total | ~$44-59 |
Prerequisites
Before starting, you need:
- An Azure VM with OpenClaw installed, an LLM connected, and Telegram working (see our Azure Setup guide)
- API keys stored properly (see our Secret Management guide)
- About 40 minutes of uninterrupted time
Step 1: Create the Google Sheet
This spreadsheet is both the input (search keywords) and output (job results) for the entire system.
Go to sheets.google.com and create a new spreadsheet. Name it something like "Job Search". Create two tabs: Criteria and Results.
Criteria tab. your search keywords (row 1 = headers):
| Column A: keyword | Column B: location |
|---|---|
| software engineer | Netherlands |
| data analyst | Netherlands |
| product manager | Netherlands |
| DevOps engineer | Netherlands |
Replace these with whatever roles and location you're searching for. Add as many rows as you need.
Results tab. where the agent writes results (row 1 = headers, leave rows empty):
| Column | Header |
|---|---|
| A | date_found |
| B | source |
| C | company |
| D | job_title |
| E | location |
| F | description |
| G | apply_url |
| H | status |
| I | language |
The status column stays empty. you fill it in manually when you apply or dismiss a listing.
Save the spreadsheet URL. The spreadsheet ID is the long string in the URL between /d/ and /edit. you'll need it later.
Step 2: Set Up Google Sheets API
Your VM cannot open a browser and log into Google. A service account is like a robot Google account that authenticates with a key file instead of a password.
- Go to console.cloud.google.com. Create a new project (name it anything, e.g. "job-search").
- Go to APIs & Services > Library. Search for "Google Sheets API". Click it, click Enable.
- Go to APIs & Services > Credentials. Click + Create Credentials > Service account. Name it, skip the optional steps.
- On the Credentials page, click your new service account. Go to the Keys tab. Click Add Key > Create new key > JSON. A
.jsonfile downloads - this is your credentials file. - Open your Google Sheet. Click Share. Paste the service account email (looks like
[email protected]). Give it Editor access. - Upload the credentials file to your VM:
# From your laptop (adjust the filename):
scp ~/Downloads/your-credentials-file.json \
openclaw@<YOUR_VM_IP>:~/.openclaw/google-credentials.json
Lock it down:
chmod 600 ~/.openclaw/google-credentials.json
Step 3: Set Up Job Search APIs
Adzuna (free)
Adzuna is a job aggregator that pulls from Indeed, company career pages, and other job boards. Their API is free for personal use.
- Go to developer.adzuna.com and create an account
- After signup, copy your Application ID and Application Key
Apify (paid. $29/month)
Apify is a cloud scraping platform. You pay them to run a scraper that collects LinkedIn job listings. No LinkedIn login needed.
- Go to apify.com and create an account
- Upgrade to the Starter plan ($29/month). the free plan doesn't have enough usage for daily scraping
- Go to Settings > Integrations and copy your Personal API Token
- Find the LinkedIn Jobs scraper at apify.com/cheap_scraper/linkedin-job-scraper
Important Apify gotchas: The
keywordfield in the API is an array (not a string). The API returns status 201 on success (not 200). Field names in the response arejobTitle,companyName, andjobUrl- nottitle,company, andurl. Add 30+ seconds delay between keyword searches or LinkedIn returns 403 errors.
Step 4: Store Credentials
Create an environment file with all your API keys. The Python scripts read this file directly.
nano ~/.openclaw/scripts/job-search.env
# Job Search Credentials
ADZUNA_APP_ID=your_adzuna_app_id
ADZUNA_APP_KEY=your_adzuna_app_key
APIFY_API_TOKEN=your_apify_token
GOOGLE_SHEETS_CREDENTIALS_PATH=/home/openclaw/.openclaw/google-credentials.json
JOB_SHEET_ID=your_google_sheet_id_here
TELEGRAM_CHAT_ID=your_telegram_chat_id
Finding your Telegram chat ID: Message
@userinfoboton Telegram - it replies instantly with your chat ID.
Lock it down:
chmod 600 ~/.openclaw/scripts/job-search.env
Step 5: Install Python Dependencies
pip install gspread google-auth requests --break-system-packages
The --break-system-packages flag is required on Ubuntu 24.04. This is safe for our use case.
Step 6: Create the Scraper Script
This is the core of the system. a Python script that calls both APIs, collects results, and pre-filters them. It runs outside OpenClaw so no LLM tokens are burned during the slow API calls.
Create the file:
nano ~/.openclaw/scripts/scrape_jobs.py
#!/usr/bin/env python3
"""
Job Search Scraper
==================
Calls Apify (LinkedIn) and Adzuna APIs, pre-filters results,
saves to JSON for the OpenClaw agent to process.
Runs OUTSIDE of OpenClaw (no LLM tokens burned while waiting).
"""
import gspread
from google.oauth2.service_account import Credentials
import requests
import json
import os
import sys
import time
from datetime import datetime, timezone
# --- Load .env file (system cron doesn't source ~/.bashrc) ---
ENV_FILE = os.path.expanduser("~/.openclaw/scripts/job-search.env")
def load_env_file(path):
if not os.path.exists(path):
return
with open(path) as f:
for line in f:
line = line.strip()
if not line or line.startswith("#"):
continue
if "=" in line:
key, value = line.split("=", 1)
os.environ.setdefault(key.strip(), value.strip())
load_env_file(ENV_FILE)
# --- Configuration ---
OUTPUT_FILE = os.path.expanduser("~/.openclaw/scripts/job_results_raw.json")
APIFY_ACTOR = "cheap_scraper~linkedin-job-scraper"
APIFY_BASE_URL = "https://api.apify.com/v2/acts"
ADZUNA_BASE_URL = "https://api.adzuna.com/v1/api/jobs/nl/search/1"
APIFY_MAX_RESULTS = 50
ADZUNA_RESULTS_PER_PAGE = 50
ADZUNA_MAX_DAYS_OLD = 7
APIFY_TIMEOUT_SECONDS = 300
APIFY_DELAY_SECONDS = 30
SCOPES = ["https://www.googleapis.com/auth/spreadsheets.readonly"]
# --- Title filter: customize these to match your target roles ---
RELEVANT_TITLE_TERMS = [
# Add terms that should appear in the job title.
# Only listings containing at least one of these terms pass the filter.
# Example for software engineering:
"software engineer", "developer", "backend", "frontend",
"full stack", "devops", "sre", "platform engineer",
"data engineer", "cloud engineer",
]
# --- Step 1: Read keywords from Google Sheet ---
def read_keywords():
print("[1/4] Reading keywords from Google Sheet...")
creds_path = os.environ.get("GOOGLE_SHEETS_CREDENTIALS_PATH")
sheet_id = os.environ.get("JOB_SHEET_ID")
if not creds_path or not sheet_id:
print(" ERROR: Missing credentials path or sheet ID")
return []
try:
creds = Credentials.from_service_account_file(creds_path, scopes=SCOPES)
client = gspread.authorize(creds)
sheet = client.open_by_key(sheet_id)
criteria = sheet.worksheet("Criteria")
rows = criteria.get_all_records()
print(f" Found {len(rows)} keywords")
return rows
except Exception as e:
print(f" ERROR: {e}")
return []
# --- Step 2: Call Apify LinkedIn scraper ---
def scrape_linkedin(keywords):
print("[2/4] Calling Apify LinkedIn scraper...")
token = os.environ.get("APIFY_API_TOKEN")
if not token:
print(" ERROR: Missing APIFY_API_TOKEN")
return [], ["Apify: missing token"]
url = f"{APIFY_BASE_URL}/{APIFY_ACTOR}/run-sync-get-dataset-items"
all_results = []
errors = []
for i, kw in enumerate(keywords):
keyword = kw.get("keyword", "")
location = kw.get("location", "Netherlands")
print(f" [{i+1}/{len(keywords)}] LinkedIn: {keyword}")
body = {
"keyword": [keyword], # Must be an array!
"location": location,
"maxResults": APIFY_MAX_RESULTS,
}
try:
resp = requests.post(url, params={"token": token},
headers={"Content-Type": "application/json"},
json=body, timeout=APIFY_TIMEOUT_SECONDS)
if resp.status_code in (200, 201):
items = resp.json()
for item in items:
all_results.append({
"source": "LinkedIn",
"job_title": item.get("jobTitle") or "",
"company": item.get("companyName") or "",
"location": item.get("location") or "",
"description": item.get("description") or "",
"apply_url": item.get("jobUrl") or "",
"keyword_used": keyword,
})
print(f" Got {len(items)} results")
else:
errors.append(f"Status {resp.status_code} for '{keyword}'")
except requests.exceptions.Timeout:
errors.append(f"Timeout for '{keyword}'")
except Exception as e:
errors.append(f"Error for '{keyword}': {e}")
if i < len(keywords) - 1:
print(f" Waiting {APIFY_DELAY_SECONDS}s...")
time.sleep(APIFY_DELAY_SECONDS)
print(f" LinkedIn total: {len(all_results)} results, {len(errors)} errors")
return all_results, errors
# --- Step 3: Call Adzuna API ---
def scrape_adzuna(keywords):
print("[3/4] Calling Adzuna API...")
app_id = os.environ.get("ADZUNA_APP_ID")
app_key = os.environ.get("ADZUNA_APP_KEY")
if not app_id or not app_key:
print(" ERROR: Missing Adzuna credentials")
return [], ["Adzuna: missing credentials"]
all_results = []
errors = []
for i, kw in enumerate(keywords):
keyword = kw.get("keyword", "")
print(f" [{i+1}/{len(keywords)}] Adzuna: {keyword}")
params = {
"app_id": app_id, "app_key": app_key,
"what": keyword,
"results_per_page": ADZUNA_RESULTS_PER_PAGE,
"max_days_old": ADZUNA_MAX_DAYS_OLD,
"content-type": "application/json",
}
try:
resp = requests.get(ADZUNA_BASE_URL, params=params, timeout=30)
if resp.status_code == 200:
items = resp.json().get("results", [])
for item in items:
all_results.append({
"source": "Adzuna",
"job_title": item.get("title") or "",
"company": (item.get("company") or {}).get("display_name") or "",
"location": (item.get("location") or {}).get("display_name") or "",
"description": item.get("description") or "",
"apply_url": item.get("redirect_url") or "",
"keyword_used": keyword,
})
print(f" Got {len(items)} results")
else:
errors.append(f"Adzuna status {resp.status_code} for '{keyword}'")
except Exception as e:
errors.append(f"Adzuna error: {e}")
print(f" Adzuna total: {len(all_results)} results, {len(errors)} errors")
return all_results, errors
# --- Step 4: Pre-filter ---
def pre_filter(results):
print("[4/4] Pre-filtering...")
# Deduplicate by URL
seen_urls = set()
unique = []
for r in results:
url = r.get("apply_url", "").strip()
if url and url not in seen_urls:
seen_urls.add(url)
unique.append(r)
# Filter by title relevance
filtered = []
for r in unique:
title = r.get("job_title", "").lower()
if any(term in title for term in RELEVANT_TITLE_TERMS):
filtered.append(r)
print(f" Raw: {len(results)} -> Deduped: {len(unique)} -> Title filter: {len(filtered)}")
return filtered
# --- Main ---
def main():
print(f"=== Job Search Scraper. {datetime.now().strftime('%Y-%m-%d %H:%M')} ===")
keywords = read_keywords()
if not keywords:
print("No keywords found. Exiting.")
sys.exit(1)
linkedin_results, linkedin_errors = scrape_linkedin(keywords)
adzuna_results, adzuna_errors = scrape_adzuna(keywords)
all_results = linkedin_results + adzuna_results
filtered = pre_filter(all_results)
output = {
"scraped_at": datetime.now(timezone.utc).isoformat(),
"keywords_used": len(keywords),
"raw_count": len(all_results),
"filtered_count": len(filtered),
"errors": linkedin_errors + adzuna_errors,
"results": filtered,
}
with open(OUTPUT_FILE, "w") as f:
json.dump(output, f, indent=2)
print(f"\nDone. {len(filtered)} results saved to {OUTPUT_FILE}")
if linkedin_errors or adzuna_errors:
print(f"Errors: {linkedin_errors + adzuna_errors}")
sys.exit(2) # Partial success
if __name__ == "__main__":
main()
Make it executable:
chmod +x ~/.openclaw/scripts/scrape_jobs.py
Customize the
RELEVANT_TITLE_TERMSlist at the top of the script. This is your title filter. only job listings containing at least one of these terms in the title will pass through. This cuts thousands of irrelevant results down to a manageable number for the AI agent.
Adzuna country code: The URL uses
/nl/for Netherlands. Change this to your country:/gb/for UK,/us/for US,/de/for Germany,/fr/for France. See Adzuna's API docs for all supported countries.
Step 7: Create the Wrapper Script
This bash script ties everything together. Cron runs this script, which runs the scraper, then triggers the OpenClaw agent to process the results.
nano ~/.openclaw/scripts/run_job_search.sh
#!/bin/bash
# Job Search Wrapper. called by cron
set -euo pipefail
SCRIPT_DIR="$HOME/.openclaw/scripts"
RESULTS_FILE="$SCRIPT_DIR/job_results_raw.json"
echo "=== Job Search Run. $(date) ==="
# Step 1: Run the scraper
python3 "$SCRIPT_DIR/scrape_jobs.py"
SCRAPER_EXIT=$?
if [ $SCRAPER_EXIT -eq 1 ]; then
echo "Scraper failed completely. Aborting."
exit 1
fi
# Step 2: Check if there are results to process
RESULT_COUNT=$(python3 -c "import json; d=json.load(open('$RESULTS_FILE')); print(d['filtered_count'])")
if [ "$RESULT_COUNT" -eq 0 ]; then
echo "No results to process. Done."
exit 0
fi
echo "Processing $RESULT_COUNT results with OpenClaw agent..."
# Step 3: Trigger the OpenClaw agent
openclaw agent \
--session isolated \
-m "Read the file $RESULTS_FILE. It contains $RESULT_COUNT pre-filtered job listings in JSON format. For each listing: 1) Check if it already exists in the Google Sheet Results tab (deduplicate by job title + company). 2) If it's new, evaluate if it's a genuine match for the search criteria. 3) Write new matches to the Results tab. 4) Send a Telegram digest with the count of new jobs and the top 5 matches (title, company, location, URL)."
echo "=== Done. $(date) ==="
Make it executable:
chmod +x ~/.openclaw/scripts/run_job_search.sh
Note: If you have a dedicated agent for job search (instead of using your main agent), add
--agent jobsearchto theopenclaw agentcommand.
Step 8: Create the Helper Scripts
The agent needs Python scripts to read from and write to your Google Sheet. Create three small scripts:
sheets_read_criteria.py
nano ~/.openclaw/scripts/sheets_read_criteria.py
#!/usr/bin/env python3
"""Read search keywords from the Criteria tab."""
import gspread, json, os, sys
from google.oauth2.service_account import Credentials
ENV_FILE = os.path.expanduser("~/.openclaw/scripts/job-search.env")
with open(ENV_FILE) as f:
for line in f:
line = line.strip()
if line and not line.startswith("#") and "=" in line:
k, v = line.split("=", 1)
os.environ.setdefault(k.strip(), v.strip())
creds = Credentials.from_service_account_file(
os.environ["GOOGLE_SHEETS_CREDENTIALS_PATH"],
scopes=["https://www.googleapis.com/auth/spreadsheets.readonly"])
client = gspread.authorize(creds)
sheet = client.open_by_key(os.environ["JOB_SHEET_ID"])
print(json.dumps(sheet.worksheet("Criteria").get_all_records(), indent=2))
sheets_read_results.py
nano ~/.openclaw/scripts/sheets_read_results.py
#!/usr/bin/env python3
"""Read existing results from the Results tab (for deduplication)."""
import gspread, json, os, sys
from google.oauth2.service_account import Credentials
ENV_FILE = os.path.expanduser("~/.openclaw/scripts/job-search.env")
with open(ENV_FILE) as f:
for line in f:
line = line.strip()
if line and not line.startswith("#") and "=" in line:
k, v = line.split("=", 1)
os.environ.setdefault(k.strip(), v.strip())
creds = Credentials.from_service_account_file(
os.environ["GOOGLE_SHEETS_CREDENTIALS_PATH"],
scopes=["https://www.googleapis.com/auth/spreadsheets.readonly"])
client = gspread.authorize(creds)
sheet = client.open_by_key(os.environ["JOB_SHEET_ID"])
print(json.dumps(sheet.worksheet("Results").get_all_records(), indent=2))
sheets_write_results.py
nano ~/.openclaw/scripts/sheets_write_results.py
#!/usr/bin/env python3
"""Append new job results to the Results tab. Reads JSON from stdin."""
import gspread, json, os, sys
from google.oauth2.service_account import Credentials
ENV_FILE = os.path.expanduser("~/.openclaw/scripts/job-search.env")
with open(ENV_FILE) as f:
for line in f:
line = line.strip()
if line and not line.startswith("#") and "=" in line:
k, v = line.split("=", 1)
os.environ.setdefault(k.strip(), v.strip())
COLUMNS = ["date_found", "source", "company", "job_title",
"location", "description", "apply_url", "status", "language"]
jobs = json.loads(sys.stdin.read())
if not jobs:
print("No jobs to write.")
sys.exit(0)
creds = Credentials.from_service_account_file(
os.environ["GOOGLE_SHEETS_CREDENTIALS_PATH"],
scopes=["https://www.googleapis.com/auth/spreadsheets"])
client = gspread.authorize(creds)
sheet = client.open_by_key(os.environ["JOB_SHEET_ID"])
results_tab = sheet.worksheet("Results")
rows = [[str(job.get(col, "")) for col in COLUMNS] for job in jobs]
results_tab.append_rows(rows, value_input_option="USER_ENTERED")
print(f"Added {len(rows)} rows to Results tab.")
Make all scripts executable:
chmod +x ~/.openclaw/scripts/*.py
Step 9: Set Up Cron
Set your VM timezone (if not already done):
sudo timedatectl set-timezone Europe/Amsterdam
Create the log directory:
mkdir -p ~/.openclaw/logs
Add the cron entries:
(crontab -l 2>/dev/null; echo "0 11 * * * $HOME/.openclaw/scripts/run_job_search.sh >> $HOME/.openclaw/logs/job-search.log 2>&1") | crontab -
(crontab -l 2>/dev/null; echo "0 15 * * * $HOME/.openclaw/scripts/run_job_search.sh >> $HOME/.openclaw/logs/job-search.log 2>&1") | crontab -
Verify:
crontab -l
You should see two entries running at 11:00 and 15:00. Change the times to whatever suits you.
Step 10: Test the Full Chain
Run the wrapper script manually:
bash ~/.openclaw/scripts/run_job_search.sh
This takes 10-20 minutes. Watch the output for keyword progress, Apify delays, pre-filter stats, and the agent processing results.
What success looks like:
- The scraper prints progress for each keyword
- Pre-filter reduces the raw count significantly
- The agent processes the filtered results
- New rows appear in your Google Sheet Results tab
- A Telegram digest arrives
Run it again to verify deduplication works. the second run should add very few new jobs because the agent compares against existing rows in the Results tab.
Costs
| Component | What you pay |
|---|---|
| Apify Starter plan | $29/month (required for daily LinkedIn scraping) |
| Adzuna API | Free |
| AI agent tokens | ~$0.25-0.50 per run, ~$15-30/month for twice daily |
| Google Sheets API | Free |
| Telegram API | Free |
The Apify plan is the biggest fixed cost. If you only need Adzuna results (no LinkedIn), you can skip Apify entirely and run for free except for LLM tokens.
Troubleshooting
| Problem | Likely cause | Fix |
|---|---|---|
| Apify returns 403 | LinkedIn rate limiting | Increase APIFY_DELAY_SECONDS in the script |
| Apify returns 0 results | Wrong input format or LinkedIn blocking | Check Apify console for error logs |
| Adzuna returns irrelevant jobs | Keywords too broad | Tighten RELEVANT_TITLE_TERMS in the script |
| Duplicate entries appearing | Dedup not matching | Normalize title comparison in the agent prompt |
| Telegram message not arriving | Wrong chat ID | Message @userinfobot to verify your chat ID |
| Google Sheet errors | Service account not shared | Re-share the sheet with the service account email |
| Cron not firing | VM deallocated or wrong timezone | Check timedatectl and crontab -l |
Check logs:
cat ~/.openclaw/logs/job-search.log | tail -50
What's Next
Once the job search runs reliably for a week:
- Tune your keywords. remove ones that produce noise, add new ones based on what you're actually interested in
- Tighten the title filter. edit
RELEVANT_TITLE_TERMSin the scraper to reduce false positives - Archive old results. move rows older than 30 days to an Archive tab in the Sheet
- Add more sources. the scraper can be extended with additional job APIs (Indeed, Glassdoor, etc.)
Check out our other guides:
- Secret Management. store your API keys properly
- Morning Briefing. automated daily briefing via Telegram
What's Next?
Before going further, make sure your API keys are properly secured. The Secret Management guide shows you how to use .env files, Azure Key Vault, and auth-profiles.json.
Just getting started with OpenClaw? The Azure Setup guide covers everything from VM creation to your first Telegram message.