Production-ready today (ship with confidence)
These nine have crossed from pilot to boring: support ticket triage and routing; tier-1 support resolution within policy (refunds, account changes); lead research and enrichment; personalized outreach drafting (human-send); CRM data hygiene (dedupe, field completion, activity logging); meeting prep briefs from calendar and CRM; internal knowledge Q&A grounded in company docs; report generation on schedule; and monitoring/alerting with plain-language summaries.
Common thread: high volume, clear policies, cheap errors or human final-send. This is where first agents should live — Chalk Labs builds most of these as 3–6 week engagements from $15k.
- Support triage + tier-1 resolution
- Lead research, enrichment, outreach drafts
- CRM hygiene and activity capture
- Meeting prep and internal knowledge Q&A
- Scheduled reporting and monitoring summaries
Ready with guardrails (ship carefully)
Eight more work in production with human checkpoints: full email send on low-stakes sequences; invoice and expense processing with approval gates; contract first-pass review flagging clauses for counsel; candidate screening summaries (bias audits mandatory); social listening with drafted responses; inventory reorder recommendations; customer onboarding orchestration; and QA test generation for engineering teams.
The pattern: judgment calls with real consequences, so the agent proposes and a human disposes — which still typically compresses the human time by 60–80%. Design the checkpoint UX as carefully as the agent; rubber-stamp fatigue is the failure mode.
Still demo-ware (wait or pilot cheaply)
Four categories that dazzle in demos and disappoint in production: fully-autonomous sales closing (negotiation across turns still fails unpredictably), open-ended research agents promising 'analyst replacement' (confident synthesis of wrong facts), autonomous financial trading/treasury actions (never, without hard limits and human keys), and multi-agent 'virtual companies' (compounding error rates across agent handoffs).
The useful reframe: today's demo-ware defines next year's roadmap. Pilot cheap, instrument honestly, and let your production agents earn the trust budget the ambitious ones will need.
Questions we hear about this
Support triage or CRM hygiene: both are high-volume, low-stakes, immediately measurable and politically easy (nobody mourns manual data entry). Success there funds and de-risks the more ambitious workflows on this list.
Eventually, but don't start there — scope creep is the leading cause of agent-project death. One workflow, instrumented, in production, then expand from evidence. Three narrow agents that work beat one broad agent that almost does.
The labor-math ones: support tier-1 resolution and CRM hygiene, where baseline human minutes per case are measurable and volume is steady. Payback in 2–6 months is typical at SMB scale against $15k–50k builds.
Demand production references for your specific use case, ask to see the eval suite (if they don't have one, run), and pilot on your real data with your real edge cases before any large commitment. Demo environments are marketing; your ticket queue is truth.