Saltar a contenido

Diagramas del sistema de chat híbrido

Verificados contra el código real el 2026-10-10. Cuando el código difiere de la descripción de la tarea, el diagrama refleja el código y lleva nota ANOTACIÓN.

1. Legacy vs híbrido + seam (ModeDispatchingChatAiResponder)

mode_dispatcher.py: stream_response resuelve el modo por conversación y delega. Hay dos fail-open a legacy, no uno: error al resolver el modo E híbrido sin responder disponible (ImportError → warning + legacy).

flowchart TD
    A["POST .../messages<br/>(conversations_router)"] --> B["ModeDispatchingChatAiResponder<br/>.stream_response"]
    B --> C["_resolve_mode<br/>conversation_id → restaurant_id"]
    C -->|conversación inexistente| L1(["LEGACY"])
    C -->|policy None o excepción| L1
    C -->|policy.chat_mode| D{mode == HYBRID?}
    D -->|no| L2(["LegacyChatAiResponder<br/>.stream_response"])
    D -->|sí| E["_build_hybrid_responder"]
    E -->|ImportError → None<br/>+ warning| L2
    E -->|HybridChatAiResponder| H(["hybrid.stream_response"])
    L1 --> L2

ANOTACIÓN: el segundo fail-open (chat_mode=hybrid pero sin responder híbrido) no estaba en la descripción; existe en el código (mode_dispatcher.py:40-48).

2. TreeInterpreter: recorrido del árbol DATA

interpreter.py: run construye ctx (state_loader.build_ctx), camina con _walk (máx. 50 pasos, interpreter_max_steps) y cierra con _finish (persiste mensaje + ChatLog + tree_state, emite done). Cualquier excepción emite un frame error.

flowchart TD
    R(["run: user_message"]) --> CTX["_build_ctx<br/>sesión, fase, carrito, tab, state"]
    CTX --> W{"_walk: nodo = entry<br/>máx 50 pasos"}
    W -->|Entry| NX["next"]
    W -->|StateCheck / Condition<br/>PhaseSwitch| BR["rama on_true / on_false<br/>branches / default"]
    W -->|Classify| CL["_exec_classify:<br/>intent, conf → branches<br/>si conf &lt; min → fallback<br/>si clasificador falla → fallback"]
    W -->|Action| AC["_exec_action<br/>timeout 5 s; dry_run salta<br/>side_effects → on_success/on_fail"]
    W -->|LlmStream| LS["_exec_llm_stream<br/>content / : thinking<br/>+ reroute a rama DATA"]
    W -->|ExecuteTags| TG["aplica cart + coursing<br/>on_changes / on_no_changes"]
    W -->|PhaseTransition<br/>PublishEvent / SetState| MUT["muta ctx / publica evento<br/>→ next"]
    NX & BR & CL & AC & LS & TG & MUT --> W
    W -->|End| FIN["_finish: persiste +<br/>recommendation? + done"]
    W -->|excepción / pasos agotados| ERR(["frame error"])

3. Árbol hybrid_v1 (simplificado)

default_trees/hybrid_v1.json (version: hybrid-1, entry check_pausa): una puerta de pausa, un classify con 6 intents y un fallback común llm_actual → aplicar_fase → ejecutar_tags → fin.

flowchart TD
    E(["entry: check_pausa<br/>state_check sesion_pausada"])
    E -->|true| LLM["llm_actual<br/>prompt_builder_por_fase"]
    E -->|false| C{"clasificar<br/>min_confidence 0.75"}
    C -->|pedir_cuenta| A1["transition_to_checkout → cuenta_llm"]
    C -->|recomendacion| A2["recommend_dishes → rec_llm"]
    C -->|respuesta_alergias| L1["alergias_ack (llm directo)"]
    C -->|consulta_alergeno| A3["lookup_allergens → alergeno_llm"]
    C -->|marchar_curso| A4["fire_course → marchar_ok/error_llm"]
    C -->|consultar_pedido| A5["get_tab → pedido_llm"]
    C -->|fallback / conf baja| LLM
    A1 & A2 & L1 & A3 & A4 & A5 --> F1
    LLM --> F2["aplicar_fase<br/>phase_transition from_tag"]
    F2 --> T["ejecutar_tags<br/>edit_cart, add_to_cart,<br/>set_course_plan, fire_course"]
    T --> F1(["fin: end"])

4. Clasificador kNN + embeddings

intent_classifier.py (k=3, blend 0.9·max + 0.1·media top-k) sobre granite_embeddings.py (singleton perezoso, vectores normalizados → producto escalar = coseno). Best-effort: sin modelo devuelve None y el llamante usa el fallback LLM. El shadow mode del legacy (engine._shadow_classify) registra intent/confidence en chat_logs en ambos motores.

flowchart LR
    MSG["mensaje usuario"] --> ENC["GraniteEmbeddingService.encode<br/>MiniLM-L12-v2, lazy, normalizado"]
    ENC --> Q["vector consulta"]
    MAT["matriz ejemplos<br/>por intent (codificada 1 vez)"] --> SIM["sims = matriz · vector<br/>(cosenos)"]
    Q --> SIM
    SIM --> KNN{"_knn_vote por intent<br/>0.9·max + 0.1·media top-3"}
    KNN --> V["(intent, confidence)<br/>o None si no hay modelo"]
    V --> U{conf ≥ min?}
    U -->|sí| BR(["rama branches[intent]"])
    U -->|no / None| FB(["fallback → llm_actual"])
    V -.->|shadow, ambos motores| LOG[("chat_logs<br/>intent_predicted/confidence")]

ANOTACIÓN (doble): el umbral pedido (0.65) es solo el default global (intent_min_confidence en config.py); el árbol real hybrid_v1.json fija min_confidence: 0.75, que prevalece. Y el "Granite" del nombre del servicio es histórico: el modelo cargado es paraphrase-multilingual-MiniLM-L12-v2 (venció a granite-97m en el benchmark, ver config.py:123-128).

5. Secuencia SSE de un turno

Contrato del stream de chat (conversations_router.send_message + intérprete): tokens content, keepalive : thinking (comentario SSE, sin data:), frame recommendation solo híbrido antes de done, y error en fallo. El router inyecta timestamp aditivo a cada frame data: (_inject_timestamp).

sequenceDiagram
    participant F as Frontend
    participant R as conversations_router
    participant I as TreeInterpreter
    F->>R: POST .../messages (SSE)
    R->>I: run/walk del árbol
    loop tokens del LLM
        I->>F: data: {"type": "content", "data": tok}
    end
    I->>F: : thinking (keepalive, comentario SSE)
    I->>F: data: {"type": "recommendation", ...} (solo híbrido, si hay recs)
    I->>F: data: {"type": "done", ...latency, model, cart/course_actions}
    Note over I,F: ante excepción: data: {"type": "error", ...} en su lugar

ANOTACIÓN: sse-events.md documenta el canal bus session:{token} (chat_message, cart_updated…), NO estos frames del stream de chat; la fuente de los frames es el intérprete + el docstring del router.

6. Vía de evolución a JEV/LAYA

El ClassifyNode es el punto de sustitución: su contrato es (intent, confidence) + branches/fallback. Un clasificador JEV o LAYA solo debe implementar classify(texto) -> (intent, confidence); el resto del árbol, intérprete y acciones quedan intactos.

flowchart TD
    subgraph INTACTO ["resto intacto"]
        T["hybrid_v1.json: ramas DATA<br/>check_pausa, acciones, llm_*"]
        W["TreeInterpreter._walk<br/>umbral min_confidence, fallback"]
        A["acciones: recommend, lookup,<br/>fire_course, get_tab..."]
    end
    subgraph SWAP ["punto de sustitución"]
        C1["ClassifyNode<br/>contrato: intent/confidence"]
        K["kNN + MiniLM (hoy)"]
        J(["JEV / LAYA (futuro)<br/>mismo contrato classify()"]
        )
    end
    T --> W --> C1
    C1 --> K
    C1 -.-> J
    K & J --> A