Skip to content

[Bug]: Server crashes (TimeoutError: RPC call to execute_model timed out) when max-num-seqs= 4, tp=4 #856

Description

@AnishPahilajani

Describe the bug

ompilation, weight/kv cache quantization).                                                                                                                                                                                       
(EngineCore pid=316931) ERROR 09-10 22:33:33 [dump_input.py:72] Dumping input data for V1 LLM engine (v0.28.0) with config: model='/models/granite-3.3-8b-instruct', speculative_config=None, tokenizer='/models/granite-3.3-8b-in
struct', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.float16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=4, pip
eline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=True, quantization=None, quantization_config=None, enforce_eager=False, enable_return_routed_expert
s=False, kv_cache_dtype=auto, device_config=cpu, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin
='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cud
agraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_
model_name=/models/granite-3.3-8b-instruct, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.STOCK_TORCH_COMPILE: 1>, 'debug_dump_path': None, 'cache_dir
': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_tok
en_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [1, 2, 4, 512], 'compile_ranges_endpoints': [512], 'inductor_compile_config': {'enable_auto_f
unctionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_
inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False,
 'fuse_allreduce_rms': False, 'enable_qk_norm_rope_fusion': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False, 'fuse_qk_norm_rope_kvcache': False}, 'max_cudagraph_capture_size': None, 'dynamic_shapes_config'
: {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op
_priority=IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, enable_jit_warmup=True, enable_bf16x3_router_gemm=False, moe_backend='auto', linear
_backend='auto'),                                                                                                                                                                                                                 
(EngineCore pid=316931) ERROR 09-10 22:33:33 [dump_input.py:79] Dumping scheduler output for model execution: SchedulerOutput(scheduled_new_reqs=[], scheduled_cached_reqs=CachedRequestData(req_ids=['cmpl-bench-74ce6b89-7-0-b9a
78042', 'cmpl-bench-74ce6b89-40-0-819235d1', 'cmpl-bench-74ce6b89-41-0-a7f6e3f1', 'cmpl-bench-74ce6b89-42-0-90eb7fd2'],resumed_req_ids=set(),new_token_ids_lens=[],all_token_ids_lens={},new_block_ids=[([284],), None, None, None
],num_computed_tokens=[8192, 447, 302, 221],num_output_tokens=[8188, 443, 294, 217]), num_scheduled_tokens={cmpl-bench-74ce6b89-7-0-b9a78042: 1, cmpl-bench-74ce6b89-42-0-90eb7fd2: 1, cmpl-bench-74ce6b89-41-0-a7f6e3f1: 1, cmpl-
bench-74ce6b89-40-0-819235d1: 1}, total_num_scheduled_tokens=4, scheduled_spec_decode_tokens={}, scheduled_encoder_inputs={}, num_common_prefix_blocks=[0], finished_req_ids=[], free_encoder_mm_hashes=[], scheduled_encoder_inpu
t_stats=null, preempted_req_ids=[], has_structured_output_requests=false, pending_structured_output_tokens=false, num_invalid_spec_tokens=null, kv_connector_metadata=null, ec_connector_metadata=null, ec_manager_metadata=null, 
new_block_ids_to_zero=null, kv_cache_block_copies=null, partial_tail_offloads=null, num_spec_tokens_to_schedule=0)                                                                                                                
(EngineCore pid=316931) ERROR 09-10 22:33:33 [dump_input.py:81] Dumping scheduler stats: SchedulerStats(num_running_reqs=4, num_waiting_reqs=0, num_skipped_waiting_reqs=0, step_counter=0, current_wave=0, kv_cache_usage=0.03613
28125, iteration_details=None, prefix_cache_stats=PrefixCacheStats(reset=False, requests=0, queries=0, hits=0, preempted_requests=0, preempted_queries=0, preempted_hits=0), connector_prefix_cache_stats=None, kv_cache_eviction_
events=[], spec_decoding_stats=None, kv_connector_stats=None, waiting_lora_adapters={}, running_lora_adapters={}, cudagraph_stats=None, perf_stats=None)   


EngineCore encountered a fatal error.                                                                                                              18:33:33 [179/1819]
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348] Traceback (most recent call last):                                                                                                                                    
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/v1/executor/multiproc_executor.py", line 430, in get_response                                         
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]     status, result = mq.dequeue(timeout=dequeue_timeout)                                                                                                              
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]                      ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^                                                                                                              
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/distributed/device_communicators/shm_broadcast.py", line 889, in dequeue                              
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]     with self.acquire_read(timeout, indefinite) as buf:                                                                                                               
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]          ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^                                                                                                                       
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]   File "/usr/lib64/python3.12/contextlib.py", line 137, in __enter__                                                                                                  
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]     return next(self.gen)                                                                                                                                             
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]            ^^^^^^^^^^^^^^                                                                                                                                             
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/distributed/device_communicators/shm_broadcast.py", line 794, in acquire_read                         
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]     self._spin_condition.wait(timeout_ms=read_timeout.timeout_ms())                                                                                                   
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]                                          ^^^^^^^^^^^^^^^^^^^^^^^^^                                                                                                    
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/distributed/device_communicators/shm_broadcast.py", line 747, in timeout_ms                           
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]     raise TimeoutError                                                                                                                                                
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348] TimeoutError                                                                                                                                                          
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]                                                                                                                                                                       
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348] The above exception was the direct cause of the following exception:                                                                                                  
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]                                                                                                                                                                       
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348] Traceback (most recent call last):                                                                                                                                    
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/v1/engine/core.py", line 1339, in run_engine_core                                                     
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]     engine_core.run_busy_loop()                                                                                                                                       
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/v1/fault_tolerance/engine_core_sentinel.py", line 179, in run_with_fault_tolerance                    
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]     busy_loop_func(self)                                                                                                                                              
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/v1/engine/core.py", line 1383, in run_busy_loop                                                       
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]     self._process_engine_step()                                                                                                                                       
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/v1/engine/core.py", line 1436, in _process_engine_step                                                
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]     outputs, model_executed = self.step_fn()                                                                                                                          
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]                               ^^^^^^^^^^^^^^                                                                                                                          
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/v1/engine/core.py", line 601, in step                                                                 
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]     model_output = future.result()                                                                                                                                    
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]                    ^^^^^^^^^^^^^^^                                                                                                                                    
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/v1/executor/multiproc_executor.py", line 99, in result                                                
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]     return super().result()                                                                                                                                           
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]            ^^^^^^^^^^^^^^^^                                                                                                                                           
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]   File "/usr/lib64/python3.12/concurrent/futures/_base.py", line 449, in result                                                                                       
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]     return self.__get_result()                                                                                                                                        
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]            ^^^^^^^^^^^^^^^^^^^                                                                                                                                        
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]   File "/usr/lib64/python3.12/concurrent/futures/_base.py", line 401, in __get_result                                                                                 
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]     raise self._exception                                                                                                                                             
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/v1/executor/multiproc_executor.py", line 103, in _wait_for_response                                   
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]     response = self.aggregate(self.get_response())                                                                                                                    
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]                               ^^^^^^^^^^^^^^^^^^^                                                                                                                     
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/v1/executor/multiproc_executor.py", line 432, in get_response                                         
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348]     raise TimeoutError(f"RPC call to {method} timed out.") from e                                                                                                     
(EngineCore pid=316931) ERROR 09-10 22:33:33 [core.py:1348] TimeoutError: RPC call to execute_model timed out.                                                                                                                    
(Worker_TP0 pid=316940) INFO 09-10 22:33:33 [multiproc_executor.py:833] Parent process exited, terminating worker queues                                                                                                          
(EngineCore pid=316931) INFO 09-10 22:33:33 [multiproc_executor.py:469] [shutdown] Executor: waiting for worker exit count=4                                                                                                      
(APIServer pid=316920) ERROR 09-10 22:33:33 [async_llm.py:744] AsyncLLM output_handler failed.                                                                                                                                    
(APIServer pid=316920) ERROR 09-10 22:33:33 [async_llm.py:744] Traceback (most recent call last):                                                                                                                                 
(APIServer pid=316920) ERROR 09-10 22:33:33 [async_llm.py:744]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/v1/engine/async_llm.py", line 690, in output_handler                                               
(APIServer pid=316920) ERROR 09-10 22:33:33 [async_llm.py:744]     outputs = await engine_core.get_output_async()                                                                                                                 
(APIServer pid=316920) ERROR 09-10 22:33:33 [async_llm.py:744]               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^                                                                                                                 
(APIServer pid=316920) ERROR 09-10 22:33:33 [async_llm.py:744]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/v1/engine/core_client.py", line 1104, in get_output_async                                          
(APIServer pid=316920) ERROR 09-10 22:33:33 [async_llm.py:744]     raise self._format_exception(outputs) from None                                                                                                                
(APIServer pid=316920) ERROR 09-10 22:33:33 [async_llm.py:744] vllm.v1.engine.exceptions.EngineDeadError: EngineCore encountered an issue. See stack trace (above) for the root cause.                                            
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496] Error in completion stream generator.                                                                                                                                
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496] Traceback (most recent call last):                                                                                                                                   
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/entrypoints/openai/completion/serving.py", line 310, in completion_stream_generator                  
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496]     async for prompt_idx, res in result_generator:                                                                                                                   
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/utils/async_utils.py", line 108, in merge_async_iterators                                            
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496]     async for item in iterator:  # type: ignore[union-attr]                                                                                                          
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/v1/engine/async_llm.py", line 607, in generate                                                       
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496]     out = q.get_nowait() or await q.get()                                                                                                                            
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496]                             ^^^^^^^^^^^^^  

.
.
.

erver pid=316920) ERROR 09-10 22:33:33 [serving.py:496]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/v1/engine/output_processor.py", line 87, in get
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496]     raise output
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/entrypoints/openai/completion/serving.py", line 310, in completion_stream_generator
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496]     async for prompt_idx, res in result_generator:
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/utils/async_utils.py", line 108, in merge_async_iterators
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496]     async for item in iterator:  # type: ignore[union-attr]
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/v1/engine/async_llm.py", line 607, in generate
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496]     out = q.get_nowait() or await q.get()
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496]                             ^^^^^^^^^^^^^
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/v1/engine/output_processor.py", line 87, in get
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496]     raise output
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/v1/engine/async_llm.py", line 690, in output_handler
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496]     outputs = await engine_core.get_output_async()
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496]               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496]   File "/home/senuser/.venv/lib64/python3.12/site-packages/vllm/v1/engine/core_client.py", line 1104, in get_output_async
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496]     raise self._format_exception(outputs) from None
(APIServer pid=316920) ERROR 09-10 22:33:33 [serving.py:496] vllm.v1.engine.exceptions.EngineDeadError: EngineCore encountered an issue. See stack trace (above) for the root cause.
(APIServer pid=316920) INFO:     127.0.0.1:44100 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=316920) INFO:     127.0.0.1:44100 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error

.
.
.

How to reproduce

Server
vllm serve /models/granite-3.3-8b-instruct --max-model-len 32768 --max-num-seqs 4 --num-gpu-blocks-override 2049 --port 8000 -tp 4

Client
vllm bench serve --backend vllm --model /models/granite-3.3-8b-instruct --dataset-name custom --dataset-path prompts.jsonl --skip-chat-template --num-prompts 100 --num-warmups 2 --max-concurrency 4 --output-len 32700 --port 8000 --save-result --save-detailed --result-dir ./results

Failed requests during benchmark run detected (capping to 10):
Error 0: Internal Server Error
Error 1: Internal Server Error
Error 2: Internal Server Error
Error 3: Internal Server Error
Error 4: Internal Server Error
Error 5: Internal Server Error
Error 6: Internal Server Error
Error 7: Internal Server Error
Error 8: Internal Server Error
Error 9: Internal Server Error
tip: install termplotlib and gnuplot to plot the metrics 
============ Serving Benchmark Result ============
Successful requests:                     43        
Failed requests:                         57        
Maximum request concurrency:             4         
Benchmark duration (s):                  6909.17   
Total input tokens:                      272       
Total generated tokens:                  35028     
Request throughput (req/s):              0.01      
Output token throughput (tok/s):         5.07      
Peak output token throughput (tok/s):    12.00     
Peak concurrent requests:                5.00      
Total token throughput (tok/s):          5.11      
---------------Time to First Token----------------
Mean TTFT (ms):                          13413.25  
Median TTFT (ms):                        2051.27   
P99 TTFT (ms):                           197156.80 
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          727.23    
Median TPOT (ms):                        735.41    
P99 TPOT (ms):                           1308.60   
---------------Inter-token Latency----------------
Mean ITL (ms):                           739.20    
Median ITL (ms):                         699.87    
P99 ITL (ms):                            925.08    
==================================================

Additional context

prompts file

prompts.jsonl.zip

Checklist

  • I have searched for similar issues.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions