Transnational online service
Chatbot analytics dashboard showing performance monitoring metrics and conversation flow data
Chatbot performance monitoring in action — session flows, response latency, and resolution rate visualised in a single dashboard view.
v3.1

Technical Reference — Chatbot Analytics

Monitoring
& Performance
Reference

This reference covers the 14 core metrics tracked by the Ualtimos Pirentex monitoring layer, how each is calculated, and what threshold ranges indicate healthy versus degraded chatbot behaviour. Practical, no fluff.

14 tracked metrics
3 alert tiers
60 s data refresh rate

Resolution Rate & Fallback Rate

Tier 1

Resolution rate measures the share of conversations where the bot delivered an answer without escalating to a human agent. Fallback rate is its inverse signal — the proportion of turns where no intent matched above the confidence threshold of 0.62. Both metrics are calculated per 1,000 sessions to keep small-volume noise from distorting trends.

Metric Healthy range Warning Critical
Resolution rate ≥ 78 % 65–77 % < 65 %
Fallback rate ≤ 8 % 9–15 % > 15 %
Resolution healthy≥ 78 %
Fallback healthy≤ 8 %
Fallback critical> 15 %

Response Latency (P50 / P95 / P99)

Tier 1

Latency is logged at 3 percentile points per session window. P50 gives you the median experience; P95 and P99 surface the tail behaviour that affects roughly 1 in 20 and 1 in 100 users respectively. A healthy P99 under 380 ms indicates the NLU pipeline is not bottlenecked during peak load.

Percentile Healthy Warning Critical
P50 ≤ 80 ms 81–140 ms > 140 ms
P95 ≤ 220 ms 221–340 ms > 340 ms
P99 ≤ 380 ms 381–600 ms > 600 ms
P50 healthy≤ 80 ms
P95 healthy≤ 220 ms
P99 critical> 600 ms

Session Depth & Containment

Tier 2

Session depth counts the number of turns before either resolution or abandonment. An average depth above 11 turns often signals intent disambiguation issues — the bot is asking clarifying questions it shouldn't need. Containment rate tracks what share of sessions stayed fully within the bot without any channel switch.

Metric Healthy Review Investigate
Avg session depth 4–8 turns 9–11 turns > 11 turns
Containment rate ≥ 82 % 70–81 % < 70 %
Avg depth healthy4–8 turns
Containment healthy≥ 82 %
Containment investigate< 70 %

Interpreting Alert Tiers

The 3-tier alert model maps metric ranges to operational actions. Tier 1 fires at thresholds that affect user experience directly. Tier 2 flags degradation patterns before they reach user-visible severity. Tier 3 is informational — it records drift over 7-day rolling windows.

T1

Immediate action

Fires when a metric crosses its critical threshold. Triggers automated incident creation and notifies the on-call group within 90 seconds. Requires acknowledgement within 15 minutes per SLA.

T2

Scheduled review

Fires at warning-range values sustained for 6 consecutive minutes. Does not page on-call — instead queues a review task for the next working session. Typical cause: intent model drift after a product update.

T3

Trend monitoring

Calculated weekly from 7-day rolling averages. Highlights gradual changes invisible in daily snapshots — for example a 1.4 % per week rise in fallback rate that would only become T2-level after 4 weeks.

Signal distribution across a typical monitored deployment

Based on aggregated session data from group monitoring sessions run through the platform, these proportions reflect what participants typically observe when reviewing a mid-scale chatbot with 4,000–9,000 monthly sessions. Individual deployments will vary.

All figures are rolling 30-day medians. Outlier spikes during deployment events are excluded from these baseline calculations.

Resolution rate
83 %
Containment
79 %
P95 latency
195 ms
Fallback rate
9 %
Avg depth
6.2 t