Google Cloud Agentic AI x Elastic

Elastic On-Call Agent:
Agentic Ops with Google Cloud

Alert-triggered triage using Cloud Run, Elastic evidence, and GitHub Issues to move from signal to root cause to agent-executed remediation.

Agentic flow

Elastic MCP CHECKING Vertex AI Gemini 2.5 Pro CHECKING Cloud Run CHECKING GitHub Issues CHECKING
Elastic alert IDLE Cloud Run endpoint IDLE Evidence collection IDLE Gemini reasoning IDLE Root cause IDLE Remediation plan IDLE Apply fix IDLE Verify health IDLE GitHub issue IDLE

Demo controls

Incident status IDLE
Service checkout-api
Severity high
Root cause -
Confidence -

Live workload health

checkout-api CHECKING

Waiting for live checkout-api status.

HTTP: - | mode: -
payment-api CHECKING

Waiting for live payment-api status.

HTTP: -
customer checkout flow CHECKING

Waiting for live end-to-end checkout flow status.

HTTP: -

Agent resolved the incident

The agent repaired checkout-api runtime mode and verified the live customer checkout flow.

Verification
  • checkout-api /healthz returned 200 after remediation.
  • checkout-api /checkout returned 200 after remediation.
  • payment-api /healthz remained healthy.
  • Customer checkout flow returned to healthy.

GitHub issue created

[resolved by agent] checkout-api failure after redis_timeout runtime mode

incident agent-remediated checkout-api elastic-evidence

Waiting for remediation result.

Open GitHub issue
Issue body preview
  • Root cause: checkout-api entered redis_timeout failure mode.
  • Action taken: agent called checkout-api /admin/repair.
  • Verification: live workload checks returned healthy.
  • Follow-up: review checkout-api runtime config guardrails.

Incident output

Click Health check, then Simulate incident.

Ask Agent

Agent answer

Ask a question after simulating an incident.