Jelou - Degradation in Conversation Processing, Routing, and Automated Workflows – Detalles del incidente

Sistemas funcionando con normalidad

Degradation in Conversation Processing, Routing, and Automated Workflows

Resuelto
Interrupción mayor
Iniciado el hace 9 díasDuró alrededor de 1 hora

Afectado

Website

Interrupción mayor de 10:45 PM a 11:45 PM

https://jelou.ai

Interrupción mayor de 10:45 PM a 11:45 PM

Platform

Interrupción mayor de 10:45 PM a 11:45 PM

Multiagent Panel

Interrupción mayor de 10:45 PM a 11:45 PM

Actualizaciones
  • Después de la muerte
    UTC
    Después de la muerte

    RCA-140826

    1. Incident Summary

    During the afternoon of August 14, a partial degradation occurred in scheduled task processing, with impact observed around 16:06 (Ecuador Time / ECT), causing delays in dependent features—primarily ticket management and scheduled message delivery. Following an initial partial recovery, the service remained under observation until a subsequent surge in request volume required scaling up the available capacity of the database supporting this processing pipeline. The capacity increase was completed at 17:45, enabling progressive recovery and subsequent normalization of task execution.

    Timeline (High Level):

    Time (ECT)

    Event

    Detail

    16:06

    Degradation observed

    Delays observed in the processing of certain scheduled tasks.

    16:45

    Partial recovery

    Processing partially improves and the service remains under observation.

    17:06

    Additional load surge

    An increase in request volume raises processing times once again.

    17:10

    Confirmed

    Customer reports are corroborated through internal telemetry and monitoring.

    17:15

    Diagnosed

    The available database capacity is identified as being exceeded.

    17:45

    Mitigated

    Database capacity expansion is completed.

    18:06

    Resolved

    Processing returns to standard operational baselines.

    Incident Window: 16:06–18:06 ECT, with partial recovery during incident management and full normalization at 18:06.

    2. Impact

    The component responsible for scheduled task execution experienced latency during the incident window, impacting dependent features, primarily ticket management and scheduled message delivery. Unrelated platform functionalities continued operating normally. Impact was limited strictly to customers whose operations required the execution of these specific tasks during the incident window.

    3. Detection

    The incident was detected via internal monitoring signals alongside incoming customer reports regarding delays in scheduled message delivery and ticket management operations. These signals facilitated rapid impact assessment and guided the technical diagnostic effort.

    4. Incident Response

    Upon detection, engineering identified that the spike in request volume had exceeded the provisioned capacity of the supporting database, leading to task queuing and elevated processing latency. As a remediation measure, available database capacity was expanded, allowing the accumulated backlog to drain and progressively restoring normal processing throughput. By 18:06, standard service performance had fully resumed.

    5. Root Cause

    The root cause was insufficient database capacity to sustain an unanticipated, significant traffic surge within the scheduled task processing layer. The sudden spike acted as the primary trigger, exhausting provisioned capacity and causing backlog buildup with increased execution latency. No application bugs or logic defects were identified in connection with this event.

    6. Resolution

    The immediate remediation was scaling up the provisioned database capacity, enabling the system to drain queued tasks and progressively restore processing throughput. Following mitigation, telemetry verified that processing latency returned to baseline metrics, and the incident was marked as resolved.

    7. Preventive Actions

    #

    Action Item

    Type

    Owner

    1

    Configure automated database capacity auto-scaling to absorb traffic surges dynamically without manual intervention.

    Infrastructure

    Infrastructure

    2

    Deploy early warning alerts on database resource and capacity utilization thresholds before degradation impacts end-user services.

    Observability

    Infrastructure

  • Resuelto
    UTC
    Resuelto
    This incident has been resolved.
  • Identificado
    UTC
    Identificado
    We are continuing to work on a fix for this incident.
  • Investigando
    UTC
    Investigando

    We are investigating an unexpected degradation affecting our conversation processing and automation services.

    As a result, you may experience delays or failures in:

    • Automatic and manual conversation routing/assignment

    • Automated workflow triggers and actions

    • Conversation lifecycle events (session expirations and state updates)

    Inbound messages from connected messaging channels continue to be safely received and retained. However, real-time routing and automated actions are temporarily delayed. Our engineering team is actively working on restoring full functionality as a high-priority incident.