Databricks Job Failure Agent with Claude Code

This runs once per failed run, so a platform with a few failures a night is a few short runs a night no matter how many jobs you schedule. Each run earns its cost by doing the reading somebody would otherwise do at nine in the morning: run output, job definition, the recent history of the same job, and a search of Linear before anything is filed. Claude Code is reliable on that kind of gather across connected tools.

Loading preview…
Free to start · guided setup

Watch it work before it's live

Run a staged conversation — no account needed. The agent handles it for real while a simulated world answers its tool calls; nothing touches real accounts, and nothing is actually sent.

nightly_revenue_rollup

Priya Ramanathanacme-prod.cloud.databricks.com / job cluster, Runtime 14.3 LTS

Run 4471028 failed at 02:14 UTC after 6 minutes on task build_revenue_facts. Error: org.apache.spark.sql.AnalysisException: [UNRESOLVED_COLUMN.WITH_SUGGESTION] A column with name `order_total_usd` cannot be resolved. Did you mean one of: [order_total, currency_code]? Line 42, pos 8.

Set up in minutes

Using this template drops you into a guided setup. It asks exactly this, nothing else:

  1. Connect Databricks

    One sign-in. The agent acts through your account, scoped to what this template uses.

  2. Connect Linear

    One sign-in. The agent acts through your account, scoped to what this template uses.

  3. Connect Slack

    One sign-in. The agent acts through your account, scoped to what this template uses.

  4. Runbook notes

    What each job is for and what breaks when it does not run - which jobs are business critical, the tables and dashboards fed by them, who owns each one, and the jobs that are allowed to fail quietly overnight.

  5. Runs on Claude Code

    Preselected for this page — connect your Claude Code account during setup, or switch to NoClick's built-in models with one click.

  6. Watch it handle a test run

    A staged conversation against a simulated world — then it’s live.

Why Claude Code for this agent

Run history decides the label

The recent runs of the same job are listed and counted before the failure is named, which is what turns a red square into first failure, flaky, or broken since a date you can point at.

Reads Databricks and Linear together

Two systems inside one run, with the search results deciding whether an issue gets filed at all. Comfortable movement across APIs is what this harness is built on.

Before you fork

Two jobs fail in the same minute with the same exception. One issue or two?

Two, because the search is on the job name together with the exception class and these are different jobs. That is deliberate. A shared root cause such as an expired credential then shows up as several issues naming the same exception, which is easier to spot and to close than one issue quietly covering three broken pipelines.

Run it with a different agent

Put Databricks Job Failure Agent to work on Claude Code

Free to start. Guided setup, a test run against staged conversations, and it's live.