Class: Raif::ModelCompletion
- Inherits:
-
ApplicationRecord
- Object
- ApplicationRecord
- Raif::ModelCompletion
- Includes:
- Concerns::BooleanTimestamp, Concerns::HasAvailableModelTools, Concerns::HasRuntimeDuration, Concerns::LlmResponseParsing, Concerns::ProviderManagedToolCalls
- Defined in:
- app/models/raif/model_completion.rb
Overview
Schema Information
Table name: raif_model_completions
id :bigint not null, primary key available_model_tools :jsonb not null cache_creation_input_tokens :integer cache_read_input_tokens :integer citations :jsonb completed_at :datetime completion_tokens :integer failed_at :datetime failure_error :string failure_reason :text failure_response_body :text failure_response_status :integer llm_model_key :string not null max_completion_tokens :integer messages :jsonb not null model_api_name :string not null output_token_cost :decimal(10, 6) prompt_token_cost :decimal(10, 6) prompt_tokens :integer raw_response :text response_array :jsonb response_finish_reason :string response_format :integer default("text"), not null response_format_parameter :string response_tool_calls :jsonb retry_count :integer default(0), not null source_type :string started_at :datetime stream_response :boolean default(FALSE), not null system_prompt :text temperature :decimal(5, 3) tool_choice :string total_cost :decimal(10, 6) total_tokens :integer created_at :datetime not null updated_at :datetime not null batch_custom_id :string raif_model_completion_batch_id :bigint response_id :string source_id :bigint
Indexes
index_raif_model_completions_on_batch_custom_id (batch_custom_id) index_raif_model_completions_on_batch_id_and_custom_id (raif_model_completion_batch_id,batch_custom_id) UNIQUE WHERE (raif_model_completion_batch_id IS NOT NULL) index_raif_model_completions_on_completed_at (completed_at) index_raif_model_completions_on_created_at (created_at) index_raif_model_completions_on_failed_at (failed_at) index_raif_model_completions_on_raif_model_completion_batch_id (raif_model_completion_batch_id) index_raif_model_completions_on_source (source_type,source_id) index_raif_model_completions_on_started_at (started_at)
Foreign Keys
fk_rails_... (raif_model_completion_batch_id => raif_model_completion_batches.id)
Constant Summary collapse
- TRUNCATED_FINISH_REASONS =
Raw provider-reported finish/stop reasons that indicate the response was cut off before completing - either at the maximum output token limit, or (on Anthropic models) because the request exhausted the model's context window (model_context_window_exceeded). The response (including any tool calls in it) is incomplete and should not be trusted.
Deliberately excluded: content-filter stops (e.g. OpenAI's "content_filter" / incomplete_details.reason "content_filter"). Those responses are also cut short, but the truncation-recovery guidance ("be more concise and retry") would be wrong for them; their partial tool calls are still rejected by argument validation.
%w[max_output_tokens length max_tokens MAX_TOKENS incomplete model_context_window_exceeded].freeze
- FAILURE_RESPONSE_BODY_MAX_CHARS =
Maximum number of characters of an upstream HTTP body we persist on failure. The body usually carries the provider's actual error reason (e.g. OpenAI/Anthropic structured error JSON), which
failure_reasoncannot fit in 255 chars. 4 KB is enough to capture realistic error payloads without bloating storage. 4_000- INFERENCE_COST_EVENT_SYNCED_COLUMNS =
Columns copied onto the inference cost event. A post-terminal change to any of them (e.g. batch results applying token counts after completed_at was already set) re-syncs the event so it stays a faithful mirror.
%w[ source_type source_id llm_model_key model_api_name prompt_tokens completion_tokens total_tokens cache_read_input_tokens cache_creation_input_tokens prompt_token_cost output_token_cost total_cost retry_count raif_model_completion_batch_id ].freeze
Constants included from Concerns::LlmResponseParsing
Concerns::LlmResponseParsing::ASCII_CONTROL_CHARS
Instance Attribute Summary collapse
-
#allow_parallel_tool_calls ⇒ Object
Request-scoped (not persisted): when true, the provider request permits the model to return multiple tool calls.
-
#anthropic_prompt_caching_enabled ⇒ Object
Returns the value of attribute anthropic_prompt_caching_enabled.
-
#bedrock_prompt_caching_enabled ⇒ Object
Returns the value of attribute bedrock_prompt_caching_enabled.
Instance Method Summary collapse
- #calculate_costs ⇒ Object
- #json_response_schema ⇒ Object
- #pending? ⇒ Boolean
- #record_failure!(exception) ⇒ Object
- #set_total_tokens ⇒ Object
-
#tool_call_summary ⇒ Object
Admin-friendly summary of the tool calls this completion requested, e.g.
- #truncated? ⇒ Boolean
Methods included from Concerns::ProviderManagedToolCalls
#provider_managed_tool_calls, #sanitized_citations
Methods included from Concerns::HasRuntimeDuration
#runtime_duration, #runtime_duration_seconds, #runtime_ended_at
Methods included from Concerns::HasAvailableModelTools
Methods included from Concerns::LlmResponseParsing
#parse_html_response, #parse_json_response, #parsed_response
Instance Attribute Details
#allow_parallel_tool_calls ⇒ Object
Request-scoped (not persisted): when true, the provider request permits the model to return multiple tool calls. Adapters that can disable parallel tool use map this onto their provider parameter. Any value other than true (including nil, the default for an instance built outside Raif::Llm#chat) is treated as false (single call).
77 78 79 |
# File 'app/models/raif/model_completion.rb', line 77 def allow_parallel_tool_calls @allow_parallel_tool_calls end |
#anthropic_prompt_caching_enabled ⇒ Object
Returns the value of attribute anthropic_prompt_caching_enabled.
70 71 72 |
# File 'app/models/raif/model_completion.rb', line 70 def anthropic_prompt_caching_enabled @anthropic_prompt_caching_enabled end |
#bedrock_prompt_caching_enabled ⇒ Object
Returns the value of attribute bedrock_prompt_caching_enabled.
70 71 72 |
# File 'app/models/raif/model_completion.rb', line 70 def bedrock_prompt_caching_enabled @bedrock_prompt_caching_enabled end |
Instance Method Details
#calculate_costs ⇒ Object
179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 |
# File 'app/models/raif/model_completion.rb', line 179 def calculate_costs # Each retry resends the same prompt, so the provider charges input tokens # for every attempt. Factor in retry_count to reflect actual billing. total_attempts = (retry_count || 0) + 1 if prompt_tokens.present? && llm_config[:input_token_cost].present? self.prompt_token_cost = calculate_prompt_token_cost(total_attempts) end if completion_tokens.present? && llm_config[:output_token_cost].present? self.output_token_cost = llm_config[:output_token_cost] * completion_tokens end if prompt_token_cost.present? || output_token_cost.present? self.total_cost = (prompt_token_cost || 0) + (output_token_cost || 0) end apply_batch_inference_discount if raif_model_completion_batch_id.present? end |
#json_response_schema ⇒ Object
171 172 173 |
# File 'app/models/raif/model_completion.rb', line 171 def json_response_schema source.json_response_schema if source&.respond_to?(:json_response_schema) end |
#pending? ⇒ Boolean
103 104 105 |
# File 'app/models/raif/model_completion.rb', line 103 def pending? started_at.nil? && completed_at.nil? && failed_at.nil? end |
#record_failure!(exception) ⇒ Object
206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 |
# File 'app/models/raif/model_completion.rb', line 206 def record_failure!(exception) self.failed_at = Time.current self.failure_error = exception.class.name self.failure_reason = exception..truncate(255) # Always clear before re-populating so a second call with a different # exception kind doesn't leave stale response metadata attached. self.failure_response_status = nil self.failure_response_body = nil # Faraday errors carry the provider's HTTP status and response body — # the latter is where the actual provider-side error reason lives. Both # are nil when the failure happened before a response was received # (DNS/connection refused/timeout). if exception.is_a?(Faraday::Error) self.failure_response_status = exception.response_status body = exception.response_body self.failure_response_body = body.to_s.first(FAILURE_RESPONSE_BODY_MAX_CHARS) if body.present? end save! end |
#set_total_tokens ⇒ Object
175 176 177 |
# File 'app/models/raif/model_completion.rb', line 175 def set_total_tokens self.total_tokens ||= completion_tokens.present? && prompt_tokens.present? ? completion_tokens + prompt_tokens : nil end |
#tool_call_summary ⇒ Object
Admin-friendly summary of the tool calls this completion requested, e.g. "5: google_search_tool (4), web_search". Combines developer-managed calls (response_tool_calls) with provider-managed calls (provider_managed_tool_calls, e.g. OpenAI/Anthropic web search). nil when no tool calls were requested.
127 128 129 130 131 132 133 134 135 |
# File 'app/models/raif/model_completion.rb', line 127 def tool_call_summary names = Array(response_tool_calls).map { |call| call["name"] } names += provider_managed_tool_calls.map { |call| call["tool_name"] } names = names.compact return if names.empty? tally = names.tally.map { |name, count| count > 1 ? "#{name} (#{count})" : name } "#{names.length}: #{tally.join(", ")}" end |
#truncated? ⇒ Boolean
119 120 121 |
# File 'app/models/raif/model_completion.rb', line 119 def truncated? TRUNCATED_FINISH_REASONS.include?(response_finish_reason) end |