Databricks Job Failure Agent with Hermes

This agent is instructed to be strictly less clever than it could be: say which task failed and since when, never say why, never name the change or the person behind it, and never put a suspected cause in the title. Copying the error out and counting the history is the entire contribution. Hermes runs that discipline with Databricks and Linear connected as tools, on an open model base.

Loading preview…
Free to start · guided setup

Watch it work before it's live

Run a staged conversation — no account needed. The agent handles it for real while a simulated world answers its tool calls; nothing touches real accounts, and nothing is actually sent.

nightly_revenue_rollup

Priya Ramanathanacme-prod.cloud.databricks.com / job cluster, Runtime 14.3 LTS

Run 4471028 failed at 02:14 UTC after 6 minutes on task build_revenue_facts. Error: org.apache.spark.sql.AnalysisException: [UNRESOLVED_COLUMN.WITH_SUGGESTION] A column with name `order_total_usd` cannot be resolved. Did you mean one of: [order_total, currency_code]? Line 42, pos 8.

Set up in minutes

Using this template drops you into a guided setup. It asks exactly this, nothing else:

  1. Connect Databricks

    One sign-in. The agent acts through your account, scoped to what this template uses.

  2. Connect Linear

    One sign-in. The agent acts through your account, scoped to what this template uses.

  3. Connect Slack

    One sign-in. The agent acts through your account, scoped to what this template uses.

  4. Runbook notes

    What each job is for and what breaks when it does not run - which jobs are business critical, the tables and dashboards fed by them, who owns each one, and the jobs that are allowed to fail quietly overnight.

  5. Runs on Hermes

    Preselected for this page — connect your Hermes account during setup, or switch to NoClick's built-in models with one click.

  6. Watch it handle a test run

    A staged conversation against a simulated world — then it’s live.

Why Hermes for this agent

No cause, no blame

The urge to diagnose from a stack trace is exactly what fills trackers with confident wrong answers. The instruction removes it and leaves the evidence for a person.

Counts are checkable

First failure, flaky or persistent arrives with the counts and dates behind it, so whoever reads the issue can verify the label instead of taking it on trust.

Before you fork

We run several hundred jobs a night. Is this expensive?

It fires once per failed run rather than once per run, so the bill follows your failure rate. A platform with three or four failures a night is three or four short runs. Jobs you have already accepted as flaky can be marked low stakes in the runbook notes, which keeps them filing an issue with no Slack traffic, and a genuinely bad night costs more by design.

Run it with a different agent

Put Databricks Job Failure Agent to work on Hermes

Free to start. Guided setup, a test run against staged conversations, and it's live.