Technical Reference — Chatbot Analytics
Monitoring
& Performance
Reference
This reference covers the 14 core metrics tracked by the Ualtimos Pirentex monitoring layer, how each is calculated, and what threshold ranges indicate healthy versus degraded chatbot behaviour. Practical, no fluff.
Resolution Rate & Fallback Rate
Tier 1Resolution rate measures the share of conversations where the bot delivered an answer without escalating to a human agent. Fallback rate is its inverse signal — the proportion of turns where no intent matched above the confidence threshold of 0.62. Both metrics are calculated per 1,000 sessions to keep small-volume noise from distorting trends.
| Metric | Healthy range | Warning | Critical |
|---|---|---|---|
| Resolution rate | ≥ 78 % | 65–77 % | < 65 % |
| Fallback rate | ≤ 8 % | 9–15 % | > 15 % |
Response Latency (P50 / P95 / P99)
Tier 1Latency is logged at 3 percentile points per session window. P50 gives you the median experience; P95 and P99 surface the tail behaviour that affects roughly 1 in 20 and 1 in 100 users respectively. A healthy P99 under 380 ms indicates the NLU pipeline is not bottlenecked during peak load.
| Percentile | Healthy | Warning | Critical |
|---|---|---|---|
| P50 | ≤ 80 ms | 81–140 ms | > 140 ms |
| P95 | ≤ 220 ms | 221–340 ms | > 340 ms |
| P99 | ≤ 380 ms | 381–600 ms | > 600 ms |
Session Depth & Containment
Tier 2Session depth counts the number of turns before either resolution or abandonment. An average depth above 11 turns often signals intent disambiguation issues — the bot is asking clarifying questions it shouldn't need. Containment rate tracks what share of sessions stayed fully within the bot without any channel switch.
| Metric | Healthy | Review | Investigate |
|---|---|---|---|
| Avg session depth | 4–8 turns | 9–11 turns | > 11 turns |
| Containment rate | ≥ 82 % | 70–81 % | < 70 % |
Interpreting Alert Tiers
The 3-tier alert model maps metric ranges to operational actions. Tier 1 fires at thresholds that affect user experience directly. Tier 2 flags degradation patterns before they reach user-visible severity. Tier 3 is informational — it records drift over 7-day rolling windows.
Immediate action
Fires when a metric crosses its critical threshold. Triggers automated incident creation and notifies the on-call group within 90 seconds. Requires acknowledgement within 15 minutes per SLA.
Scheduled review
Fires at warning-range values sustained for 6 consecutive minutes. Does not page on-call — instead queues a review task for the next working session. Typical cause: intent model drift after a product update.
Trend monitoring
Calculated weekly from 7-day rolling averages. Highlights gradual changes invisible in daily snapshots — for example a 1.4 % per week rise in fallback rate that would only become T2-level after 4 weeks.
Signal distribution across a typical monitored deployment
Based on aggregated session data from group monitoring sessions run through the platform, these proportions reflect what participants typically observe when reviewing a mid-scale chatbot with 4,000–9,000 monthly sessions. Individual deployments will vary.
All figures are rolling 30-day medians. Outlier spikes during deployment events are excluded from these baseline calculations.