What You'll Build

An automated job search pipeline that runs twice a day on your Azure VM. It scrapes LinkedIn and Adzuna for job listings matching your keywords, filters them with AI, tracks results in a Google Sheet, and sends a Telegram digest.

The flow:

  1. Linux cron fires at the times you choose (e.g. 11 AM and 3 PM)
  2. A Python scraper calls Apify (LinkedIn) and Adzuna (job boards) for all your keywords
  3. The scraper pre-filters results: removes duplicates, checks title relevance
  4. The scraper saves everything to a JSON file on disk
  5. A bash wrapper triggers your OpenClaw agent
  6. The agent reads the JSON, does AI-powered filtering, deduplicates against your Google Sheet, writes new jobs, and sends a Telegram digest

Why this architecture? The scraper runs outside OpenClaw because API calls take 10-20 minutes. Running them inside the agent would burn tokens while waiting. The agent only sees pre-filtered results, so it processes them in seconds instead of minutes.

Estimated monthly cost:

Component Monthly
Apify LinkedIn scraper (Starter plan) $29
Adzuna API Free
AI filtering tokens (Claude or GPT) ~$15-30
Google Sheets API + Telegram Free
Total ~$44-59

Prerequisites

Before starting, you need:

Step 1: Create the Google Sheet

This spreadsheet is both the input (search keywords) and output (job results) for the entire system.

Go to sheets.google.com and create a new spreadsheet. Name it something like "Job Search". Create two tabs: Criteria and Results.

Criteria tab. your search keywords (row 1 = headers):

Column A: keyword Column B: location
software engineer Netherlands
data analyst Netherlands
product manager Netherlands
DevOps engineer Netherlands

Replace these with whatever roles and location you're searching for. Add as many rows as you need.

Results tab. where the agent writes results (row 1 = headers, leave rows empty):

Column Header
A date_found
B source
C company
D job_title
E location
F description
G apply_url
H status
I language

The status column stays empty. you fill it in manually when you apply or dismiss a listing.

Save the spreadsheet URL. The spreadsheet ID is the long string in the URL between /d/ and /edit. you'll need it later.

Step 2: Set Up Google Sheets API

Your VM cannot open a browser and log into Google. A service account is like a robot Google account that authenticates with a key file instead of a password.

  1. Go to console.cloud.google.com. Create a new project (name it anything, e.g. "job-search").
  2. Go to APIs & Services > Library. Search for "Google Sheets API". Click it, click Enable.
  3. Go to APIs & Services > Credentials. Click + Create Credentials > Service account. Name it, skip the optional steps.
  4. On the Credentials page, click your new service account. Go to the Keys tab. Click Add Key > Create new key > JSON. A .json file downloads - this is your credentials file.
  5. Open your Google Sheet. Click Share. Paste the service account email (looks like [email protected]). Give it Editor access.
  6. Upload the credentials file to your VM:
# From your laptop (adjust the filename):
scp ~/Downloads/your-credentials-file.json \
  openclaw@<YOUR_VM_IP>:~/.openclaw/google-credentials.json

Lock it down:

chmod 600 ~/.openclaw/google-credentials.json

Step 3: Set Up Job Search APIs

Adzuna (free)

Adzuna is a job aggregator that pulls from Indeed, company career pages, and other job boards. Their API is free for personal use.

  1. Go to developer.adzuna.com and create an account
  2. After signup, copy your Application ID and Application Key

Apify (paid. $29/month)

Apify is a cloud scraping platform. You pay them to run a scraper that collects LinkedIn job listings. No LinkedIn login needed.

  1. Go to apify.com and create an account
  2. Upgrade to the Starter plan ($29/month). the free plan doesn't have enough usage for daily scraping
  3. Go to Settings > Integrations and copy your Personal API Token
  4. Find the LinkedIn Jobs scraper at apify.com/cheap_scraper/linkedin-job-scraper

Important Apify gotchas: The keyword field in the API is an array (not a string). The API returns status 201 on success (not 200). Field names in the response are jobTitle, companyName, and jobUrl - not title, company, and url. Add 30+ seconds delay between keyword searches or LinkedIn returns 403 errors.

Step 4: Store Credentials

Create an environment file with all your API keys. The Python scripts read this file directly.

nano ~/.openclaw/scripts/job-search.env
# Job Search Credentials
ADZUNA_APP_ID=your_adzuna_app_id
ADZUNA_APP_KEY=your_adzuna_app_key
APIFY_API_TOKEN=your_apify_token
GOOGLE_SHEETS_CREDENTIALS_PATH=/home/openclaw/.openclaw/google-credentials.json
JOB_SHEET_ID=your_google_sheet_id_here
TELEGRAM_CHAT_ID=your_telegram_chat_id

Finding your Telegram chat ID: Message @userinfobot on Telegram - it replies instantly with your chat ID.

Lock it down:

chmod 600 ~/.openclaw/scripts/job-search.env

Step 5: Install Python Dependencies

pip install gspread google-auth requests --break-system-packages

The --break-system-packages flag is required on Ubuntu 24.04. This is safe for our use case.

Step 6: Create the Scraper Script

This is the core of the system. a Python script that calls both APIs, collects results, and pre-filters them. It runs outside OpenClaw so no LLM tokens are burned during the slow API calls.

Create the file:

nano ~/.openclaw/scripts/scrape_jobs.py
#!/usr/bin/env python3
"""
Job Search Scraper
==================
Calls Apify (LinkedIn) and Adzuna APIs, pre-filters results,
saves to JSON for the OpenClaw agent to process.

Runs OUTSIDE of OpenClaw (no LLM tokens burned while waiting).
"""

import gspread
from google.oauth2.service_account import Credentials
import requests
import json
import os
import sys
import time
from datetime import datetime, timezone

# --- Load .env file (system cron doesn't source ~/.bashrc) ---
ENV_FILE = os.path.expanduser("~/.openclaw/scripts/job-search.env")

def load_env_file(path):
    if not os.path.exists(path):
        return
    with open(path) as f:
        for line in f:
            line = line.strip()
            if not line or line.startswith("#"):
                continue
            if "=" in line:
                key, value = line.split("=", 1)
                os.environ.setdefault(key.strip(), value.strip())

load_env_file(ENV_FILE)

# --- Configuration ---
OUTPUT_FILE = os.path.expanduser("~/.openclaw/scripts/job_results_raw.json")
APIFY_ACTOR = "cheap_scraper~linkedin-job-scraper"
APIFY_BASE_URL = "https://api.apify.com/v2/acts"
ADZUNA_BASE_URL = "https://api.adzuna.com/v1/api/jobs/nl/search/1"
APIFY_MAX_RESULTS = 50
ADZUNA_RESULTS_PER_PAGE = 50
ADZUNA_MAX_DAYS_OLD = 7
APIFY_TIMEOUT_SECONDS = 300
APIFY_DELAY_SECONDS = 30
SCOPES = ["https://www.googleapis.com/auth/spreadsheets.readonly"]

# --- Title filter: customize these to match your target roles ---
RELEVANT_TITLE_TERMS = [
    # Add terms that should appear in the job title.
    # Only listings containing at least one of these terms pass the filter.
    # Example for software engineering:
    "software engineer", "developer", "backend", "frontend",
    "full stack", "devops", "sre", "platform engineer",
    "data engineer", "cloud engineer",
]

# --- Step 1: Read keywords from Google Sheet ---
def read_keywords():
    print("[1/4] Reading keywords from Google Sheet...")
    creds_path = os.environ.get("GOOGLE_SHEETS_CREDENTIALS_PATH")
    sheet_id = os.environ.get("JOB_SHEET_ID")
    if not creds_path or not sheet_id:
        print("  ERROR: Missing credentials path or sheet ID")
        return []
    try:
        creds = Credentials.from_service_account_file(creds_path, scopes=SCOPES)
        client = gspread.authorize(creds)
        sheet = client.open_by_key(sheet_id)
        criteria = sheet.worksheet("Criteria")
        rows = criteria.get_all_records()
        print(f"  Found {len(rows)} keywords")
        return rows
    except Exception as e:
        print(f"  ERROR: {e}")
        return []

# --- Step 2: Call Apify LinkedIn scraper ---
def scrape_linkedin(keywords):
    print("[2/4] Calling Apify LinkedIn scraper...")
    token = os.environ.get("APIFY_API_TOKEN")
    if not token:
        print("  ERROR: Missing APIFY_API_TOKEN")
        return [], ["Apify: missing token"]

    url = f"{APIFY_BASE_URL}/{APIFY_ACTOR}/run-sync-get-dataset-items"
    all_results = []
    errors = []

    for i, kw in enumerate(keywords):
        keyword = kw.get("keyword", "")
        location = kw.get("location", "Netherlands")
        print(f"  [{i+1}/{len(keywords)}] LinkedIn: {keyword}")

        body = {
            "keyword": [keyword],  # Must be an array!
            "location": location,
            "maxResults": APIFY_MAX_RESULTS,
        }

        try:
            resp = requests.post(url, params={"token": token},
                                 headers={"Content-Type": "application/json"},
                                 json=body, timeout=APIFY_TIMEOUT_SECONDS)
            if resp.status_code in (200, 201):
                items = resp.json()
                for item in items:
                    all_results.append({
                        "source": "LinkedIn",
                        "job_title": item.get("jobTitle") or "",
                        "company": item.get("companyName") or "",
                        "location": item.get("location") or "",
                        "description": item.get("description") or "",
                        "apply_url": item.get("jobUrl") or "",
                        "keyword_used": keyword,
                    })
                print(f"    Got {len(items)} results")
            else:
                errors.append(f"Status {resp.status_code} for '{keyword}'")
        except requests.exceptions.Timeout:
            errors.append(f"Timeout for '{keyword}'")
        except Exception as e:
            errors.append(f"Error for '{keyword}': {e}")

        if i < len(keywords) - 1:
            print(f"    Waiting {APIFY_DELAY_SECONDS}s...")
            time.sleep(APIFY_DELAY_SECONDS)

    print(f"  LinkedIn total: {len(all_results)} results, {len(errors)} errors")
    return all_results, errors

# --- Step 3: Call Adzuna API ---
def scrape_adzuna(keywords):
    print("[3/4] Calling Adzuna API...")
    app_id = os.environ.get("ADZUNA_APP_ID")
    app_key = os.environ.get("ADZUNA_APP_KEY")
    if not app_id or not app_key:
        print("  ERROR: Missing Adzuna credentials")
        return [], ["Adzuna: missing credentials"]

    all_results = []
    errors = []

    for i, kw in enumerate(keywords):
        keyword = kw.get("keyword", "")
        print(f"  [{i+1}/{len(keywords)}] Adzuna: {keyword}")

        params = {
            "app_id": app_id, "app_key": app_key,
            "what": keyword,
            "results_per_page": ADZUNA_RESULTS_PER_PAGE,
            "max_days_old": ADZUNA_MAX_DAYS_OLD,
            "content-type": "application/json",
        }

        try:
            resp = requests.get(ADZUNA_BASE_URL, params=params, timeout=30)
            if resp.status_code == 200:
                items = resp.json().get("results", [])
                for item in items:
                    all_results.append({
                        "source": "Adzuna",
                        "job_title": item.get("title") or "",
                        "company": (item.get("company") or {}).get("display_name") or "",
                        "location": (item.get("location") or {}).get("display_name") or "",
                        "description": item.get("description") or "",
                        "apply_url": item.get("redirect_url") or "",
                        "keyword_used": keyword,
                    })
                print(f"    Got {len(items)} results")
            else:
                errors.append(f"Adzuna status {resp.status_code} for '{keyword}'")
        except Exception as e:
            errors.append(f"Adzuna error: {e}")

    print(f"  Adzuna total: {len(all_results)} results, {len(errors)} errors")
    return all_results, errors

# --- Step 4: Pre-filter ---
def pre_filter(results):
    print("[4/4] Pre-filtering...")

    # Deduplicate by URL
    seen_urls = set()
    unique = []
    for r in results:
        url = r.get("apply_url", "").strip()
        if url and url not in seen_urls:
            seen_urls.add(url)
            unique.append(r)

    # Filter by title relevance
    filtered = []
    for r in unique:
        title = r.get("job_title", "").lower()
        if any(term in title for term in RELEVANT_TITLE_TERMS):
            filtered.append(r)

    print(f"  Raw: {len(results)} -> Deduped: {len(unique)} -> Title filter: {len(filtered)}")
    return filtered

# --- Main ---
def main():
    print(f"=== Job Search Scraper. {datetime.now().strftime('%Y-%m-%d %H:%M')} ===")

    keywords = read_keywords()
    if not keywords:
        print("No keywords found. Exiting.")
        sys.exit(1)

    linkedin_results, linkedin_errors = scrape_linkedin(keywords)
    adzuna_results, adzuna_errors = scrape_adzuna(keywords)

    all_results = linkedin_results + adzuna_results
    filtered = pre_filter(all_results)

    output = {
        "scraped_at": datetime.now(timezone.utc).isoformat(),
        "keywords_used": len(keywords),
        "raw_count": len(all_results),
        "filtered_count": len(filtered),
        "errors": linkedin_errors + adzuna_errors,
        "results": filtered,
    }

    with open(OUTPUT_FILE, "w") as f:
        json.dump(output, f, indent=2)

    print(f"\nDone. {len(filtered)} results saved to {OUTPUT_FILE}")

    if linkedin_errors or adzuna_errors:
        print(f"Errors: {linkedin_errors + adzuna_errors}")
        sys.exit(2)  # Partial success

if __name__ == "__main__":
    main()

Make it executable:

chmod +x ~/.openclaw/scripts/scrape_jobs.py

Customize the RELEVANT_TITLE_TERMS list at the top of the script. This is your title filter. only job listings containing at least one of these terms in the title will pass through. This cuts thousands of irrelevant results down to a manageable number for the AI agent.

Adzuna country code: The URL uses /nl/ for Netherlands. Change this to your country: /gb/ for UK, /us/ for US, /de/ for Germany, /fr/ for France. See Adzuna's API docs for all supported countries.

Step 7: Create the Wrapper Script

This bash script ties everything together. Cron runs this script, which runs the scraper, then triggers the OpenClaw agent to process the results.

nano ~/.openclaw/scripts/run_job_search.sh
#!/bin/bash
# Job Search Wrapper. called by cron
set -euo pipefail

SCRIPT_DIR="$HOME/.openclaw/scripts"
RESULTS_FILE="$SCRIPT_DIR/job_results_raw.json"

echo "=== Job Search Run. $(date) ==="

# Step 1: Run the scraper
python3 "$SCRIPT_DIR/scrape_jobs.py"
SCRAPER_EXIT=$?

if [ $SCRAPER_EXIT -eq 1 ]; then
    echo "Scraper failed completely. Aborting."
    exit 1
fi

# Step 2: Check if there are results to process
RESULT_COUNT=$(python3 -c "import json; d=json.load(open('$RESULTS_FILE')); print(d['filtered_count'])")

if [ "$RESULT_COUNT" -eq 0 ]; then
    echo "No results to process. Done."
    exit 0
fi

echo "Processing $RESULT_COUNT results with OpenClaw agent..."

# Step 3: Trigger the OpenClaw agent
openclaw agent \
  --session isolated \
  -m "Read the file $RESULTS_FILE. It contains $RESULT_COUNT pre-filtered job listings in JSON format. For each listing: 1) Check if it already exists in the Google Sheet Results tab (deduplicate by job title + company). 2) If it's new, evaluate if it's a genuine match for the search criteria. 3) Write new matches to the Results tab. 4) Send a Telegram digest with the count of new jobs and the top 5 matches (title, company, location, URL)."

echo "=== Done. $(date) ==="

Make it executable:

chmod +x ~/.openclaw/scripts/run_job_search.sh

Note: If you have a dedicated agent for job search (instead of using your main agent), add --agent jobsearch to the openclaw agent command.

Step 8: Create the Helper Scripts

The agent needs Python scripts to read from and write to your Google Sheet. Create three small scripts:

sheets_read_criteria.py

nano ~/.openclaw/scripts/sheets_read_criteria.py
#!/usr/bin/env python3
"""Read search keywords from the Criteria tab."""
import gspread, json, os, sys
from google.oauth2.service_account import Credentials

ENV_FILE = os.path.expanduser("~/.openclaw/scripts/job-search.env")
with open(ENV_FILE) as f:
    for line in f:
        line = line.strip()
        if line and not line.startswith("#") and "=" in line:
            k, v = line.split("=", 1)
            os.environ.setdefault(k.strip(), v.strip())

creds = Credentials.from_service_account_file(
    os.environ["GOOGLE_SHEETS_CREDENTIALS_PATH"],
    scopes=["https://www.googleapis.com/auth/spreadsheets.readonly"])
client = gspread.authorize(creds)
sheet = client.open_by_key(os.environ["JOB_SHEET_ID"])
print(json.dumps(sheet.worksheet("Criteria").get_all_records(), indent=2))

sheets_read_results.py

nano ~/.openclaw/scripts/sheets_read_results.py
#!/usr/bin/env python3
"""Read existing results from the Results tab (for deduplication)."""
import gspread, json, os, sys
from google.oauth2.service_account import Credentials

ENV_FILE = os.path.expanduser("~/.openclaw/scripts/job-search.env")
with open(ENV_FILE) as f:
    for line in f:
        line = line.strip()
        if line and not line.startswith("#") and "=" in line:
            k, v = line.split("=", 1)
            os.environ.setdefault(k.strip(), v.strip())

creds = Credentials.from_service_account_file(
    os.environ["GOOGLE_SHEETS_CREDENTIALS_PATH"],
    scopes=["https://www.googleapis.com/auth/spreadsheets.readonly"])
client = gspread.authorize(creds)
sheet = client.open_by_key(os.environ["JOB_SHEET_ID"])
print(json.dumps(sheet.worksheet("Results").get_all_records(), indent=2))

sheets_write_results.py

nano ~/.openclaw/scripts/sheets_write_results.py
#!/usr/bin/env python3
"""Append new job results to the Results tab. Reads JSON from stdin."""
import gspread, json, os, sys
from google.oauth2.service_account import Credentials

ENV_FILE = os.path.expanduser("~/.openclaw/scripts/job-search.env")
with open(ENV_FILE) as f:
    for line in f:
        line = line.strip()
        if line and not line.startswith("#") and "=" in line:
            k, v = line.split("=", 1)
            os.environ.setdefault(k.strip(), v.strip())

COLUMNS = ["date_found", "source", "company", "job_title",
           "location", "description", "apply_url", "status", "language"]

jobs = json.loads(sys.stdin.read())
if not jobs:
    print("No jobs to write.")
    sys.exit(0)

creds = Credentials.from_service_account_file(
    os.environ["GOOGLE_SHEETS_CREDENTIALS_PATH"],
    scopes=["https://www.googleapis.com/auth/spreadsheets"])
client = gspread.authorize(creds)
sheet = client.open_by_key(os.environ["JOB_SHEET_ID"])
results_tab = sheet.worksheet("Results")
rows = [[str(job.get(col, "")) for col in COLUMNS] for job in jobs]
results_tab.append_rows(rows, value_input_option="USER_ENTERED")
print(f"Added {len(rows)} rows to Results tab.")

Make all scripts executable:

chmod +x ~/.openclaw/scripts/*.py

Step 9: Set Up Cron

Set your VM timezone (if not already done):

sudo timedatectl set-timezone Europe/Amsterdam

Create the log directory:

mkdir -p ~/.openclaw/logs

Add the cron entries:

(crontab -l 2>/dev/null; echo "0 11 * * * $HOME/.openclaw/scripts/run_job_search.sh >> $HOME/.openclaw/logs/job-search.log 2>&1") | crontab -
(crontab -l 2>/dev/null; echo "0 15 * * * $HOME/.openclaw/scripts/run_job_search.sh >> $HOME/.openclaw/logs/job-search.log 2>&1") | crontab -

Verify:

crontab -l

You should see two entries running at 11:00 and 15:00. Change the times to whatever suits you.

Step 10: Test the Full Chain

Run the wrapper script manually:

bash ~/.openclaw/scripts/run_job_search.sh

This takes 10-20 minutes. Watch the output for keyword progress, Apify delays, pre-filter stats, and the agent processing results.

What success looks like:

Run it again to verify deduplication works. the second run should add very few new jobs because the agent compares against existing rows in the Results tab.

Costs

Component What you pay
Apify Starter plan $29/month (required for daily LinkedIn scraping)
Adzuna API Free
AI agent tokens ~$0.25-0.50 per run, ~$15-30/month for twice daily
Google Sheets API Free
Telegram API Free

The Apify plan is the biggest fixed cost. If you only need Adzuna results (no LinkedIn), you can skip Apify entirely and run for free except for LLM tokens.

Troubleshooting

Problem Likely cause Fix
Apify returns 403 LinkedIn rate limiting Increase APIFY_DELAY_SECONDS in the script
Apify returns 0 results Wrong input format or LinkedIn blocking Check Apify console for error logs
Adzuna returns irrelevant jobs Keywords too broad Tighten RELEVANT_TITLE_TERMS in the script
Duplicate entries appearing Dedup not matching Normalize title comparison in the agent prompt
Telegram message not arriving Wrong chat ID Message @userinfobot to verify your chat ID
Google Sheet errors Service account not shared Re-share the sheet with the service account email
Cron not firing VM deallocated or wrong timezone Check timedatectl and crontab -l

Check logs:

cat ~/.openclaw/logs/job-search.log | tail -50

What's Next

Once the job search runs reliably for a week:

Check out our other guides:

What's Next?

Before going further, make sure your API keys are properly secured. The Secret Management guide shows you how to use .env files, Azure Key Vault, and auth-profiles.json.

Just getting started with OpenClaw? The Azure Setup guide covers everything from VM creation to your first Telegram message.