deerflow2

History

Xinmin Zeng be0eae9825 fix(runtime): suppress tool execution when provider safety-terminates with tool_calls (#3035 ) * fix(runtime): suppress tool execution when provider safety-terminates with tool_calls When a provider stops generation for safety reasons (OpenAI/Moonshot finish_reason=content_filter, Anthropic stop_reason=refusal, Gemini finish_reason=SAFETY/BLOCKLIST/PROHIBITED_CONTENT/SPII/RECITATION/ IMAGE_SAFETY/...), the response may still carry truncated tool_calls. LangChain's tool router treats any non-empty tool_calls as executable, so partial arguments (e.g. write_file with a half-finished markdown) get dispatched and the agent loops on retry. Add SafetyFinishReasonMiddleware at after_model: detect safety termination via a pluggable detector registry, clear both structured tool_calls and raw additional_kwargs.tool_calls / function_call, preserve response_metadata.finish_reason for downstream observers, stamp additional_kwargs.safety_termination for traces, append a user-facing explanation to message content (list-aware for thinking blocks), and emit a safety_termination custom stream event so SSE consumers can reconcile any "tool starting..." UI. Default detectors cover OpenAI-compatible content_filter, Anthropic refusal, and Gemini safety enums (text + image). Custom providers are added via reflection (same pattern as guardrails). Wired into both lead-agent and subagent runtimes. Closes #3028 * fix(runtime): persist safety_termination as a middleware audit event Address review on #3035: the SSE custom event is great for live consumers but invisible to post-run audit. RunEventStore should carry its own row so operators can answer "which runs were safety-suppressed today?" from a single SQL query without joining the message body. Worker now exposes the run-scoped RunJournal via runtime.context["__run_journal"] (sentinel key, internal channel). SafetyFinishReasonMiddleware calls the previously-unused RunJournal.record_middleware, which emits event_type = "middleware:safety_termination" category = "middleware" content = {name, hook, action, changes={ detector, reason_field, reason_value, suppressed_tool_call_count, suppressed_tool_call_names, suppressed_tool_call_ids, message_id, extras}} Tool arguments are deliberately excluded — those are the very content the provider filtered and persisting them would defeat the purpose of the safety filter (per review note in #3035). Graceful skips when journal is absent (subagent runtime, unit tests, no-event-store local dev). Journal exceptions never propagate into the agent loop. Refs #3028 * fix(runtime): satisfy ruff format + address Copilot review - ruff format on safety_finish_reason_config.py and e2e demo (CI lint failed on ruff format --check; backend Makefile lint target runs ruff check AND ruff format --check). - Docstring on SafetyFinishReasonConfig now says resolve_variable to match the actual loader used in from_config (the wording was resolve_class previously; behavior is unchanged — resolve_variable mirrors how guardrails.provider is loaded). - Switch the AIMessage type check in SafetyFinishReasonMiddleware._apply from getattr(last, "type") == "ai" to isinstance(last, AIMessage), matching TokenUsageMiddleware / TodoMiddleware / ViewImageMiddleware / SummarizationMiddleware which are the dominant pattern. Refs #3028		2026-05-22 21:20:28 +08:00
..
__init__.py	feat: add create_deerflow_agent SDK entry point (Phase 1) (#1203 )	2026-03-29 15:31:18 +08:00
clarification_middleware.py	fix(backend): make clarification messages idempotent (#2350 ) (#2351 )	2026-04-19 22:00:58 +08:00
dangling_tool_call_middleware.py	fix(middleware): normalize tool result adjacency before model calls (#2939 )	2026-05-15 22:09:04 +08:00
deferred_tool_filter_middleware.py	fix: gate deferred MCP tool execution (#2513 )	2026-04-24 22:45:41 +08:00
dynamic_context_middleware.py	fix(lint): remove duplicate is_dynamic_context_reminder definition (#2837 )	2026-05-09 23:40:46 +08:00
llm_error_handling_middleware.py	refactor: thread app_config through middleware factories (#2652 )	2026-04-30 12:41:09 +08:00
loop_detection_middleware.py	feat(loop-detection): defer warning injection (#2752 )	2026-05-21 14:36:07 +08:00
memory_middleware.py	refactor: thread app_config through lead and subagent task path (#2666 )	2026-05-02 06:37:49 +08:00
safety_finish_reason_middleware.py	fix(runtime): suppress tool execution when provider safety-terminates with tool_calls (#3035 )	2026-05-22 21:20:28 +08:00
safety_termination_detectors.py	fix(runtime): suppress tool execution when provider safety-terminates with tool_calls (#3035 )	2026-05-22 21:20:28 +08:00
sandbox_audit_middleware.py	feat(sandbox): strengthen bash command auditing with compound splitting and expanded patterns (#1881 )	2026-04-07 17:15:24 +08:00
subagent_limit_middleware.py	fix(middleware): sync raw tool call metadata (#2757 )	2026-05-08 10:08:53 +08:00
summarization_middleware.py	fix(harness): preserve dynamic context across summarization (#2823 )	2026-05-09 19:39:36 +08:00
thread_data_middleware.py	feat: enhance chat history loading with new hooks and UI components (#2338 )	2026-04-26 11:20:17 +08:00
title_middleware.py	fix(tracing): propagate session_id and user_id into Langfuse traces (#2944 )	2026-05-21 16:49:31 +08:00
todo_middleware.py	fix(middleware): Prevent todo completion reminder IMMessage leak (#2907 )	2026-05-15 22:12:37 +08:00
token_usage_middleware.py	feat: stream subagent token usage to header via terminal task events (#2882 )	2026-05-13 23:52:19 +08:00
tool_call_metadata.py	fix(middleware): sync raw tool call metadata (#2757 )	2026-05-08 10:08:53 +08:00
tool_error_handling_middleware.py	fix(runtime): suppress tool execution when provider safety-terminates with tool_calls (#3035 )	2026-05-22 21:20:28 +08:00
uploads_middleware.py	feat: enhance chat history loading with new hooks and UI components (#2338 )	2026-04-26 11:20:17 +08:00
view_image_middleware.py	fix(backend): preserve viewed image reducer metadata (#1900 )	2026-04-06 16:47:19 +08:00