diff --git a/CHANGELOG.md b/CHANGELOG.md index c1bea144..bfdfe40c 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,28 @@ +# Python SDK Changelog + ## [Unreleased] +## [1.2.0] - 2026-06-04 + +### Added + +- **Metric versions API: `client.metric_versions`** + - New `client.metric_versions.list(metric_id)`, `.create(metric_id, request)`, and `.deploy(metric_id, version_name)` methods (plus `*_async` variants) for managing a metric's immutable version history, backed by `/v1/metrics/{metric_id}/versions`. Each metric keeps a history of versions with one deployed at a time; pass `deploy_immediately=True` on create to atomically deploy the new version. Version request/response models are exported from `honeyhive.models`. + +### Fixed + +- **OTLP JSON exporter: serialize numeric fields as JSON strings** + - Integer span attributes (`intValue`) and uint64 timestamp fields + (`startTimeUnixNano`, `endTimeUnixNano`, `timeUnixNano`) were emitted as + raw JSON numbers. The protobuf JSON mapping spec requires these to be JSON + strings; values above 2^53 could otherwise lose precision silently through + float64 rounding. Now matches native `opentelemetry-exporter-otlp-proto-http` + behavior. +- **Decorator input capture: remove truncation of serialized list/dict values** + - List and dict span attributes are serialized to strings and + previously were truncated to 1000 characters in length. The arbitrary limit + of 1000 characters has been removed. + ## [1.1.0] - 2026-05-19 ### Added @@ -49,7 +72,7 @@ ### Removed - **`honeyhive` Python CLI entry point removed from `pyproject.toml`** - - The shipped Python `honeyhive` console script was non-functional (dead code) and shadowed the official TypeScript CLI on `$PATH`. CLI functionality is now provided by the official [`honeyhive` TypeScript CLI](https://docs.honeyhive.ai/v2/sdk-reference/cli); removing the Python script entry point lets `honeyhive` resolve correctly when both packages are installed globally. + - The shipped Python `honeyhive` console script was non-functional (dead code) and shadowed the official TypeScript CLI on `$PATH`. CLI functionality is now provided by the official [`honeyhive` TypeScript CLI](https://docs.honeyhive.ai/v2/cli-reference/getting-started); removing the Python script entry point lets `honeyhive` resolve correctly when both packages are installed globally. ## [1.0.1] - 2026-05-05 @@ -82,6 +105,8 @@ The changes below are relative to `1.0.0rc22`. For the full picture of what ship - **`DeleteMetricQuery` removed from `honeyhive.models`** - Internal query-params type that was exported but unused — no public method accepted or returned it. The public `client.metrics.delete(id=...)` signature is unchanged. + + ## [1.0.0rc22] - 2026-05-01 ### Added diff --git a/openapi/v1.yaml b/openapi/v1.yaml index 371a5cd8..59b11434 100644 --- a/openapi/v1.yaml +++ b/openapi/v1.yaml @@ -66,7 +66,7 @@ paths: $ref: '#/components/schemas/SessionTracesResponse' '400': description: | - Bad request — both `logs` and `events` supplied, or neither, or invalid event payload. + Bad request: both `logs` and `events` supplied, or neither, or invalid event payload. x-honeyhive-plane: dp /v1/sessions: post: @@ -81,35 +81,35 @@ paths: `session` wrapper). The server creates a session event and returns it. - **No required properties** — every field has a server-side fallback. + **No required properties.** Every field has a server-side fallback. **Auto-generated properties** (provided by the server when omitted): - - `session_id` (string, UUID) — Server generates a UUIDv4 if omitted + - `session_id` (string, UUID): Server generates a UUIDv4 if omitted or if the supplied value is not a valid UUID. **Optional properties with defaults:** - - `event_name` (string) — Falls back to `session_name` when not + - `event_name` (string): Falls back to `session_name` when not provided; defaults to `"unknown"` if both are absent. - - `source` (string) — Defaults to `"unknown"`. + - `source` (string): Defaults to `"unknown"`. **Optional properties:** - - `session_name` (string) — Display name for the session. - - `start_time` (number) — Session start time as Unix milliseconds. + - `session_name` (string): Display name for the session. + - `start_time` (number): Session start time as Unix milliseconds. The session normalizer uses `getInt64()` which only accepts numeric types; if a string is passed, the server silently falls back to the current time. - - `end_time` (number) — Session end time as Unix milliseconds (same + - `end_time` (number): Session end time as Unix milliseconds (same numeric-only caveat as `start_time`). - - `duration` (number) — Session duration in milliseconds. - - `config` (object) — Configuration associated with the session. - - `inputs` (object) — Input data for the session. - - `outputs` (object) — Output data from the session. - - `metadata` (object) — Arbitrary metadata. - - `user_properties` (object) — User properties. - - `children_ids` (array of strings) — IDs of child events. + - `duration` (number): Session duration in milliseconds. + - `config` (object): Configuration associated with the session. + - `inputs` (object): Input data for the session. + - `outputs` (object): Output data from the session. + - `metadata` (object): Arbitrary metadata. + - `user_properties` (object): User properties. + - `children_ids` (array of strings): IDs of child events. Idempotent on `session_id`: posting twice with the same `session_id` merges metadata/user_properties into the existing session and returns @@ -139,23 +139,15 @@ paths: x-ts-sdk-name: createEventBatch summary: Add a batch of events to a session description: | - AIP-233 nested batch create. Adds a batch of events to an existing - session. Each event in the batch is stored with `session_id` set from - the URL path, overriding any `session_id` in the event body. + Add a batch of events to an existing session. Each event in the batch + is stored with `session_id` set from the URL path, overriding any + `session_id` in the event body. - **Required properties:** - - - `events` (array of event objects) — Each event must include - `event_type` (one of `chain`, `model`, `tool`, `session`) and `inputs`. - - Unknown top-level fields and unknown per-event fields are rejected at - the SDK boundary; the deprecated per-event `project` field is no - longer accepted. - - Events are processed sequentially (not via the worker-pool batch path - used by `POST /v1/events/batch`) — semantics match the legacy - `POST /session/{session_id}/traces` route per the Normalize Routes - RFC. + Each event must include `event_type` (one of `chain`, `model`, `tool`, + `session`) and `inputs`. Unknown top-level fields and unknown per-event + fields are rejected. Events are processed sequentially. For + higher-throughput ingestion across sessions, use + `POST /v1/events/batch` instead. parameters: - name: session_id in: path @@ -272,33 +264,33 @@ paths: **Required properties:** - - `event_type` (string) — Must be one of: `chain`, `model`, `tool`, `session`. - - `inputs` (object) — Input data for the event. + - `event_type` (string): Must be one of: `chain`, `model`, `tool`, `session`. + - `inputs` (object): Input data for the event. **Auto-generated properties** (provided by the server when omitted): - - `event_id` (string, UUID) — Unique identifier for the event. - - `session_id` (string, UUID) — Session/trace identifier. - - `parent_id` (string, UUID) — Parent event ID. Defaults to `session_id`. + - `event_id` (string, UUID): Unique identifier for the event. + - `session_id` (string, UUID): Session/trace identifier. + - `parent_id` (string, UUID): Parent event ID. Defaults to `session_id`. **Optional properties with defaults:** - - `event_name` (string) — Name of the event. Defaults to `"unknown"`. - - `source` (string) — Source of the event (e.g. `sdk-python`). Defaults to `"unknown"`. + - `event_name` (string): Name of the event. Defaults to `"unknown"`. + - `source` (string): Source of the event (e.g. `sdk-python`). Defaults to `"unknown"`. **Optional properties:** - - `config` (object) — Configuration data (e.g. model parameters, prompt templates). - - `outputs` (object) — Output data from the event. - - `error` (string or null) — Error message if the event failed. - - `children_ids` (array of strings) — IDs of child events. - - `duration` (number) — Duration of the event in milliseconds. - - `start_time` (number) — Unix timestamp in milliseconds for event start. - - `end_time` (number) — Unix timestamp in milliseconds for event end. - - `metadata` (object) — Additional metadata (e.g. token counts, cost). - - `metrics` (object) — Custom metrics. - - `feedback` (object) — Feedback data (e.g. ratings, ground truth). - - `user_properties` (object) — User properties associated with the event. + - `config` (object): Configuration data (e.g. model parameters, prompt templates). + - `outputs` (object): Output data from the event. + - `error` (string or null): Error message if the event failed. + - `children_ids` (array of strings): IDs of child events. + - `duration` (number): Duration of the event in milliseconds. + - `start_time` (number): Unix timestamp in milliseconds for event start. + - `end_time` (number): Unix timestamp in milliseconds for event end. + - `metadata` (object): Additional metadata (e.g. token counts, cost). + - `metrics` (object): Custom metrics. + - `feedback` (object): Feedback data (e.g. ratings, ground truth). + - `user_properties` (object): User properties associated with the event. requestBody: required: true content: @@ -434,14 +426,14 @@ paths: **Required properties:** - - `events` (array of event objects) — Each event must include + - `events` (array of event objects): Each event must include `event_type` (one of `chain`, `model`, `tool`, `session`) and `inputs`. **Optional properties:** - - `single_session` (boolean) — If true, all events share a single session + - `single_session` (boolean): If true, all events share a single session created from `session_properties`. Defaults to false. - - `session_properties` (object) — Session metadata used when + - `session_properties` (object): Session metadata used when `single_session` is true. May include `session_name`, `start_time`, `metadata`. @@ -543,7 +535,7 @@ paths: deprecated: true summary: Create a batch of model events (deprecated) description: | - Deprecated. Use `POST /v1/events/batch` with `event_type="model"` on each event instead. Migration notes: the top-level array `model_events` becomes `events`; each element must explicitly set `event_type: "model"` (the legacy route sets this server-side); the deprecated top-level aliases `is_single_session` and `session` are not accepted by the v1 route (`PostEventBatchRequest` is `.strict()`) — use `single_session` and `session_properties` instead. The legacy route continues to serve traffic and remaps the model-specific fields on each event (`model`, `messages`, `response`, `provider`, `usage`, `cost`, `hyperparameters`, `template`, `template_inputs`, `tools`, `tool_choice`, `response_format`, `duration`, `error`) into `inputs.*` / `outputs.*` before storage. + Deprecated. Use `POST /v1/events/batch` with `event_type="model"` on each event instead. Migration notes: the top-level array `model_events` becomes `events`; each element must explicitly set `event_type: "model"` (the legacy route sets this server-side); the deprecated top-level aliases `is_single_session` and `session` are not accepted by the v1 route (`PostEventBatchRequest` is `.strict()`); use `single_session` and `session_properties` instead. The legacy route continues to serve traffic and remaps the model-specific fields on each event (`model`, `messages`, `response`, `provider`, `usage`, `cost`, `hyperparameters`, `template`, `template_inputs`, `tools`, `tool_choice`, `response_format`, `duration`, `error`) into `inputs.*` / `outputs.*` before storage. requestBody: required: true content: @@ -589,7 +581,6 @@ paths: x-cli-name: create x-ts-sdk-name: create summary: Create a new chart - description: Add a new chart requestBody: required: true content: @@ -664,7 +655,6 @@ paths: x-cli-name: delete x-ts-sdk-name: delete summary: Delete a chart - description: Remove a chart by id. parameters: - in: path name: chart_id @@ -687,8 +677,7 @@ paths: operationId: getMetrics x-cli-name: list x-ts-sdk-name: list - summary: Get all metrics - description: Retrieve a list of all metrics + summary: List all metrics parameters: - name: type in: query @@ -717,7 +706,6 @@ paths: x-cli-name: create x-ts-sdk-name: create summary: Create a new metric - description: Add a new metric requestBody: required: true content: @@ -816,7 +804,6 @@ paths: x-cli-name: delete x-ts-sdk-name: delete summary: Delete a metric - description: Remove a metric by id. parameters: - in: path name: metric_id @@ -832,6 +819,97 @@ paths: schema: $ref: '#/components/schemas/DeleteMetricResponse' x-honeyhive-plane: dp + /v1/metrics/{metric_id}/versions: + get: + tags: + - Metric Versions + operationId: getMetricVersions + x-cli-name: list + x-ts-sdk-name: list + summary: List versions for a metric + description: Retrieve all snapshot versions of the metric's definition, ordered oldest-first. Returns the full version history unpaginated. + parameters: + - in: path + name: metric_id + required: true + schema: + type: string + description: The unique identifier of the metric whose versions are being listed + responses: + '200': + description: Metric versions retrieved successfully + content: + application/json: + schema: + $ref: '#/components/schemas/GetMetricVersionsResponse' + '404': + description: No metric with the given `metric_id` exists in the caller's scope. + x-honeyhive-plane: dp + post: + tags: + - Metric Versions + operationId: createMetricVersion + x-cli-name: create + x-ts-sdk-name: create + summary: Create a new metric version + description: 'Snapshot the supplied metric definition as a new version. By default the version is created as a draft (`deployed: false`); set `deploy_immediately: true` to also make it the live version in the same transaction.' + parameters: + - in: path + name: metric_id + required: true + schema: + type: string + description: The unique identifier of the metric to version + requestBody: + required: true + content: + application/json: + schema: + $ref: '#/components/schemas/CreateMetricVersionRequest' + responses: + '200': + description: Metric version created successfully + content: + application/json: + schema: + $ref: '#/components/schemas/CreateMetricVersionResponse' + '400': + description: Invalid `content` payload. Fails cross-field validation (e.g. an `LLM` metric without `model_provider`/`model_name`, or a `COMPOSITE` metric without `child_metrics`). + '404': + description: No metric with the given `metric_id` exists in the caller's scope. + x-honeyhive-plane: dp + /v1/metrics/{metric_id}/versions/{version_name}/deploy: + post: + tags: + - Metric Versions + operationId: deployMetricVersion + x-cli-name: deploy + x-ts-sdk-name: deploy + summary: Deploy a specific metric version + description: Mark the named version as the live version for the metric, unmarking any previously deployed version. + parameters: + - in: path + name: metric_id + required: true + schema: + type: string + description: The unique identifier of the metric + - in: path + name: version_name + required: true + schema: + type: string + description: The name of the version to deploy + responses: + '200': + description: Metric version deployed successfully + content: + application/json: + schema: + $ref: '#/components/schemas/DeployMetricVersionResponse' + '404': + description: No metric with the given `metric_id` exists in the caller's scope, or no version with the given `version_name` exists for that metric. + x-honeyhive-plane: dp /v1/metrics/run: post: tags: @@ -2244,17 +2322,22 @@ components: name: type: string maxLength: 200 + description: Display name for the chart description: type: string + description: Description of what the chart shows metric: type: string minLength: 1 + description: Name of the metric to visualize func: type: string + description: Aggregation function to apply (e.g. sum, avg, median, min, max) groupBy: type: - string - 'null' + description: Field to group results by bucketing: type: string enum: @@ -2264,17 +2347,21 @@ components: - week - month default: day + description: Time bucket granularity for aggregation dateRange: anyOf: - $ref: '#/components/schemas/RelativeDateRange' - $ref: '#/components/schemas/AbsoluteDateRange' + description: Time range to query query: type: array items: $ref: '#/components/schemas/QueryFilter' + description: Filters to apply to the chart data owner_id: type: string minLength: 1 + description: ID of the user who owns this chart required: - name - metric @@ -2306,14 +2393,18 @@ components: properties: field: type: string + description: Name of the field to filter on value: type: - string - 'null' + description: Value to compare against type: type: string + description: Data type of the field (e.g. string, number) operator: type: string + description: Comparison operator (e.g. is, is not, contains, greater than, less than) required: - field - value @@ -2372,16 +2463,21 @@ components: name: type: string maxLength: 200 + description: Display name for the chart description: type: string + description: Description of what the chart shows metric: type: string + description: Name of the metric to visualize func: type: string + description: Aggregation function to apply (e.g. sum, avg, median, min, max) groupBy: type: - string - 'null' + description: Field to group results by bucketing: type: string enum: @@ -2390,17 +2486,21 @@ components: - day - week - month + description: Time bucket granularity for aggregation dateRange: anyOf: - $ref: '#/components/schemas/RelativeDateRange' - $ref: '#/components/schemas/AbsoluteDateRange' + description: Time range to query query: type: array items: $ref: '#/components/schemas/QueryFilter' + description: Filters to apply to the chart data owner_id: type: string minLength: 1 + description: ID of the user who owns this chart additionalProperties: false UpdateChartResponse: type: object @@ -3713,7 +3813,7 @@ components: description: Page number of results (default 1) ignore_order: type: boolean - description: 'Deprecated: accepted for SDK back-compat but treated as a no-op. Pagination requires a stable ORDER BY to produce consistent pages, and with the 1000-row cap skipping the sort is not worth the inconsistency. The route always orders by start_time DESC.' + description: 'Deprecated: accepted but ignored. Results are always ordered by start_time descending.' evaluation_id: type: string description: Filter by evaluation/experiment run ID @@ -5252,6 +5352,321 @@ components: - result - explanation description: Response for POST /metrics/run + MetricVersionContent: + type: object + properties: + name: + type: string + maxLength: 200 + type: + type: string + enum: + - PYTHON + - LLM + - HUMAN + - COMPOSITE + criteria: + type: string + minLength: 1 + description: + type: + - string + - 'null' + return_type: + type: string + enum: + - float + - boolean + - string + - categorical + enabled_in_prod: + type: boolean + needs_ground_truth: + type: boolean + sampling_percentage: + type: number + minimum: 0 + maximum: 100 + model_provider: + type: + - string + - 'null' + model_name: + type: + - string + - 'null' + scale: + type: + - integer + - 'null' + exclusiveMinimum: 0 + threshold: + type: + - object + - 'null' + properties: + min: + type: number + max: + type: number + pass_when: + anyOf: + - type: boolean + - type: number + passing_categories: + type: array + items: + type: string + maxLength: 200 + minItems: 1 + additionalProperties: false + categories: + type: + - array + - 'null' + items: + $ref: '#/components/schemas/MetricVersionContentCategoriesItem' + minItems: 1 + child_metrics: + type: + - array + - 'null' + items: + $ref: '#/components/schemas/MetricVersionContentChildMetricsItem' + minItems: 1 + filters: + $ref: '#/components/schemas/MetricVersionContentFilters' + required: + - name + - type + - criteria + - description + - return_type + - enabled_in_prod + - needs_ground_truth + - sampling_percentage + - filters + additionalProperties: false + description: Versioned snapshot of a metric definition. + MetricVersionContentRequest: + type: object + properties: + name: + type: string + maxLength: 200 + type: + type: string + enum: + - PYTHON + - LLM + - HUMAN + - COMPOSITE + criteria: + type: string + minLength: 1 + description: + type: string + default: '' + description: Free-form description of the metric. Defaults to an empty string. + return_type: + type: string + enum: + - float + - boolean + - string + - categorical + default: float + description: Return type of the metric. Defaults to `float`. + enabled_in_prod: + type: boolean + default: false + description: Whether this version should run against production traffic. Defaults to `false` for non-HUMAN metrics and `true` for HUMAN metrics. + needs_ground_truth: + type: boolean + default: false + description: Whether this metric requires ground-truth labels to evaluate. Defaults to `false`. + sampling_percentage: + type: number + minimum: 0 + maximum: 100 + default: 10 + description: Percentage of events the metric should run against, 0–100. Defaults to `10`. + model_provider: + type: + - string + - 'null' + model_name: + type: + - string + - 'null' + scale: + type: + - integer + - 'null' + exclusiveMinimum: 0 + threshold: + type: + - object + - 'null' + properties: + min: + type: number + max: + type: number + pass_when: + anyOf: + - type: boolean + - type: number + passing_categories: + type: array + items: + type: string + maxLength: 200 + minItems: 1 + additionalProperties: false + categories: + type: + - array + - 'null' + items: + $ref: '#/components/schemas/MetricVersionContentRequestCategoriesItem' + minItems: 1 + child_metrics: + type: + - array + - 'null' + items: + $ref: '#/components/schemas/MetricVersionContentRequestChildMetricsItem' + minItems: 1 + filters: + $ref: '#/components/schemas/MetricVersionContentRequestFilters' + required: + - name + - type + - criteria + additionalProperties: false + description: |- + Metric definition snapshot accepted by POST /v1/metrics/{metric_id}/versions. + Six fields are optional and fall back to server-side defaults when omitted: + - `description` → `""` + - `return_type` → `"float"` + - `enabled_in_prod` → `true` for HUMAN metrics, `false` otherwise + - `needs_ground_truth` → `false` + - `sampling_percentage` → `10` + - `filters` → `{ "filterArray": [] }` + MetricVersion: + type: object + properties: + name: + type: string + full_sha: + type: string + message: + type: string + date: + type: string + format: date-time + deployed: + type: boolean + content: + $ref: '#/components/schemas/MetricVersionContent' + required: + - name + - full_sha + - message + - date + - deployed + - content + additionalProperties: false + description: A versioned snapshot of a metric definition. + GetMetricVersionsParams: + type: object + properties: + metric_id: + type: string + required: + - metric_id + additionalProperties: false + description: Path parameters for GET /v1/metrics/{metric_id}/versions + CreateMetricVersionParams: + type: object + properties: + metric_id: + type: string + required: + - metric_id + additionalProperties: false + description: Path parameters for POST /v1/metrics/{metric_id}/versions + DeployMetricVersionParams: + type: object + properties: + metric_id: + type: string + version_name: + type: string + required: + - metric_id + - version_name + additionalProperties: false + description: Path parameters for POST /v1/metrics/{metric_id}/versions/{version_name}/deploy + CreateMetricVersionRequest: + type: object + properties: + message: + type: string + content: + $ref: '#/components/schemas/MetricVersionContentRequest' + deploy_immediately: + type: boolean + required: + - message + - content + additionalProperties: false + description: Request body for POST /v1/metrics/{metric_id}/versions + GetMetricVersionsResponse: + type: object + properties: + success: + type: boolean + enum: + - true + data: + type: array + items: + $ref: '#/components/schemas/MetricVersion' + required: + - success + - data + additionalProperties: false + description: Response for GET /v1/metrics/{metric_id}/versions + CreateMetricVersionResponse: + type: object + properties: + success: + type: boolean + enum: + - true + data: + $ref: '#/components/schemas/MetricVersion' + required: + - success + - data + additionalProperties: false + description: Response for POST /v1/metrics/{metric_id}/versions + DeployMetricVersionResponse: + type: object + properties: + success: + type: boolean + enum: + - true + data: + $ref: '#/components/schemas/MetricVersion' + required: + - success + - data + additionalProperties: false + description: Response for POST /v1/metrics/{metric_id}/versions/{version_name}/deploy ProjectItem: type: object properties: @@ -6697,6 +7112,93 @@ components: feedback: $ref: '#/components/schemas/LegacyRunMetricRequestEventFeedback' additionalProperties: {} + MetricVersionContentCategoriesItem: + type: object + properties: + category: + type: string + maxLength: 200 + score: + type: + - number + - 'null' + required: + - category + - score + additionalProperties: false + MetricVersionContentChildMetricsItem: + type: object + properties: + id: + type: string + minLength: 1 + name: + type: string + maxLength: 200 + weight: + type: number + scale: + type: + - integer + - 'null' + exclusiveMinimum: 0 + required: + - name + - weight + additionalProperties: false + MetricVersionContentFilters: + type: object + properties: + filterArray: + $ref: '#/components/schemas/FiltersArray' + required: + - filterArray + additionalProperties: false + MetricVersionContentRequestCategoriesItem: + type: object + properties: + category: + type: string + maxLength: 200 + score: + type: + - number + - 'null' + required: + - category + - score + additionalProperties: false + MetricVersionContentRequestChildMetricsItem: + type: object + properties: + id: + type: string + minLength: 1 + name: + type: string + maxLength: 200 + weight: + type: number + scale: + type: + - integer + - 'null' + exclusiveMinimum: 0 + required: + - name + - weight + additionalProperties: false + MetricVersionContentRequestFilters: + type: object + properties: + filterArray: + $ref: '#/components/schemas/FiltersArray' + default: + filterArray: [] + required: + - filterArray + additionalProperties: false + description: 'ETL filter narrowing which events this metric applies to. Defaults to `{ filterArray: [] }` (no filtering).' AnnotationQueueFilters: type: object properties: @@ -7015,10 +7517,10 @@ security: tags: - name: Charts description: | - Define and manage saved charts — visualizations that aggregate metrics over time with bucketing, filters, and groupings. + Define and manage saved charts. Charts are visualizations that aggregate metrics over time with bucketing, filters, and groupings. - name: Configurations description: | - Manage prompt configurations — model parameters, message templates, and response settings — that can be versioned and deployed without code changes. + Manage prompt configurations, i.e. model parameters, message templates, and response settings that can be versioned and deployed without code changes. - name: Datapoints description: | Manage individual records inside datasets, including batch creation and mapping to source events. @@ -7027,13 +7529,16 @@ tags: Curate collections of datapoints used as test sets for evaluations and experiments. - name: Events description: | - Read and write trace events — the spans that capture every step of an AI application's execution. + Read and write trace events. Events are the spans that capture every step of an AI application's execution. - name: Experiments description: | Run, retrieve, and compare evaluation runs to measure how prompt or configuration changes affect agent performance. + - name: Metric Versions + description: | + Snapshot, list, and deploy versions of a metric's definition so changes can be reviewed and rolled back without losing history. - name: Metrics description: | - Define and run evaluators — automated quality checks that score traces against criteria like accuracy, safety, or correctness. + Define and run evaluators, i.e. automated quality checks that score traces against criteria like accuracy, safety, or correctness. - name: Queues description: | Manage annotation queues for human review of traces, turning expert feedback into labeled datasets. diff --git a/pyproject.toml b/pyproject.toml index e327b83a..43c35070 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -45,6 +45,7 @@ dev = [ "pytest-asyncio>=0.21.0", "pytest-cov>=7.0.0", "pytest-mock>=3.10.0", + "pytest-rerunfailures>=16.3", "pytest-xdist>=3.0.0", "tox>=4.0.0", "ruff>=0.4.0", diff --git a/src/honeyhive/__init__.py b/src/honeyhive/__init__.py index ff550a24..ef23d109 100644 --- a/src/honeyhive/__init__.py +++ b/src/honeyhive/__init__.py @@ -5,7 +5,7 @@ # Version must be defined BEFORE imports to avoid circular import issues # Version must be semver or semver followed by "a" (alpha), "b" (beta), or "rc" # (release candidate) + a number -__version__ = "1.1.0" +__version__ = "1.2.0" # Main API client from .api import HoneyHive diff --git a/src/honeyhive/_generated/models/CreateMetricVersionParams.py b/src/honeyhive/_generated/models/CreateMetricVersionParams.py new file mode 100644 index 00000000..4b8291b7 --- /dev/null +++ b/src/honeyhive/_generated/models/CreateMetricVersionParams.py @@ -0,0 +1,21 @@ +from typing import Any, Dict, List, Optional, Union + +from pydantic import BaseModel, Field + +__all__ = ["CreateMetricVersionParams"] + + +class CreateMetricVersionParams(BaseModel): + """ + CreateMetricVersionParams model + Path parameters for POST /v1/metrics/{metric_id}/versions + """ + + model_config = { + "populate_by_name": True, + "validate_assignment": True, + "extra": "allow", + "protected_namespaces": (), + } + + metric_id: str = Field(validation_alias="metric_id") diff --git a/src/honeyhive/_generated/models/CreateMetricVersionRequest.py b/src/honeyhive/_generated/models/CreateMetricVersionRequest.py new file mode 100644 index 00000000..e11dd89e --- /dev/null +++ b/src/honeyhive/_generated/models/CreateMetricVersionRequest.py @@ -0,0 +1,29 @@ +from typing import Any, Dict, List, Optional, Union + +from pydantic import BaseModel, Field + +from .MetricVersionContentRequest import MetricVersionContentRequest + +__all__ = ["CreateMetricVersionRequest"] + + +class CreateMetricVersionRequest(BaseModel): + """ + CreateMetricVersionRequest model + Request body for POST /v1/metrics/{metric_id}/versions + """ + + model_config = { + "populate_by_name": True, + "validate_assignment": True, + "extra": "allow", + "protected_namespaces": (), + } + + message: str = Field(validation_alias="message") + + content: MetricVersionContentRequest = Field(validation_alias="content") + + deploy_immediately: Optional[bool] = Field( + validation_alias="deploy_immediately", default=None + ) diff --git a/src/honeyhive/_generated/models/CreateMetricVersionResponse.py b/src/honeyhive/_generated/models/CreateMetricVersionResponse.py new file mode 100644 index 00000000..9b59c29f --- /dev/null +++ b/src/honeyhive/_generated/models/CreateMetricVersionResponse.py @@ -0,0 +1,25 @@ +from typing import Any, Dict, List, Optional, Union + +from pydantic import BaseModel, Field + +from .MetricVersion import MetricVersion + +__all__ = ["CreateMetricVersionResponse"] + + +class CreateMetricVersionResponse(BaseModel): + """ + CreateMetricVersionResponse model + Response for POST /v1/metrics/{metric_id}/versions + """ + + model_config = { + "populate_by_name": True, + "validate_assignment": True, + "extra": "allow", + "protected_namespaces": (), + } + + success: bool = Field(validation_alias="success") + + data: MetricVersion = Field(validation_alias="data") diff --git a/src/honeyhive/_generated/models/DeployMetricVersionParams.py b/src/honeyhive/_generated/models/DeployMetricVersionParams.py new file mode 100644 index 00000000..13a4ed85 --- /dev/null +++ b/src/honeyhive/_generated/models/DeployMetricVersionParams.py @@ -0,0 +1,23 @@ +from typing import Any, Dict, List, Optional, Union + +from pydantic import BaseModel, Field + +__all__ = ["DeployMetricVersionParams"] + + +class DeployMetricVersionParams(BaseModel): + """ + DeployMetricVersionParams model + Path parameters for POST /v1/metrics/{metric_id}/versions/{version_name}/deploy + """ + + model_config = { + "populate_by_name": True, + "validate_assignment": True, + "extra": "allow", + "protected_namespaces": (), + } + + metric_id: str = Field(validation_alias="metric_id") + + version_name: str = Field(validation_alias="version_name") diff --git a/src/honeyhive/_generated/models/DeployMetricVersionResponse.py b/src/honeyhive/_generated/models/DeployMetricVersionResponse.py new file mode 100644 index 00000000..a9f862a6 --- /dev/null +++ b/src/honeyhive/_generated/models/DeployMetricVersionResponse.py @@ -0,0 +1,25 @@ +from typing import Any, Dict, List, Optional, Union + +from pydantic import BaseModel, Field + +from .MetricVersion import MetricVersion + +__all__ = ["DeployMetricVersionResponse"] + + +class DeployMetricVersionResponse(BaseModel): + """ + DeployMetricVersionResponse model + Response for POST /v1/metrics/{metric_id}/versions/{version_name}/deploy + """ + + model_config = { + "populate_by_name": True, + "validate_assignment": True, + "extra": "allow", + "protected_namespaces": (), + } + + success: bool = Field(validation_alias="success") + + data: MetricVersion = Field(validation_alias="data") diff --git a/src/honeyhive/_generated/models/GetMetricVersionsParams.py b/src/honeyhive/_generated/models/GetMetricVersionsParams.py new file mode 100644 index 00000000..08b285e5 --- /dev/null +++ b/src/honeyhive/_generated/models/GetMetricVersionsParams.py @@ -0,0 +1,21 @@ +from typing import Any, Dict, List, Optional, Union + +from pydantic import BaseModel, Field + +__all__ = ["GetMetricVersionsParams"] + + +class GetMetricVersionsParams(BaseModel): + """ + GetMetricVersionsParams model + Path parameters for GET /v1/metrics/{metric_id}/versions + """ + + model_config = { + "populate_by_name": True, + "validate_assignment": True, + "extra": "allow", + "protected_namespaces": (), + } + + metric_id: str = Field(validation_alias="metric_id") diff --git a/src/honeyhive/_generated/models/GetMetricVersionsResponse.py b/src/honeyhive/_generated/models/GetMetricVersionsResponse.py new file mode 100644 index 00000000..8ad1fd99 --- /dev/null +++ b/src/honeyhive/_generated/models/GetMetricVersionsResponse.py @@ -0,0 +1,25 @@ +from typing import Any, Dict, List, Optional, Union + +from pydantic import BaseModel, Field + +from .MetricVersion import MetricVersion + +__all__ = ["GetMetricVersionsResponse"] + + +class GetMetricVersionsResponse(BaseModel): + """ + GetMetricVersionsResponse model + Response for GET /v1/metrics/{metric_id}/versions + """ + + model_config = { + "populate_by_name": True, + "validate_assignment": True, + "extra": "allow", + "protected_namespaces": (), + } + + success: bool = Field(validation_alias="success") + + data: List[MetricVersion] = Field(validation_alias="data") diff --git a/src/honeyhive/_generated/models/MetricVersion.py b/src/honeyhive/_generated/models/MetricVersion.py new file mode 100644 index 00000000..1669b72b --- /dev/null +++ b/src/honeyhive/_generated/models/MetricVersion.py @@ -0,0 +1,33 @@ +from typing import Any, Dict, List, Optional, Union + +from pydantic import BaseModel, Field + +from .MetricVersionContent import MetricVersionContent + +__all__ = ["MetricVersion"] + + +class MetricVersion(BaseModel): + """ + MetricVersion model + A versioned snapshot of a metric definition. + """ + + model_config = { + "populate_by_name": True, + "validate_assignment": True, + "extra": "allow", + "protected_namespaces": (), + } + + name: str = Field(validation_alias="name") + + full_sha: str = Field(validation_alias="full_sha") + + message: str = Field(validation_alias="message") + + date: str = Field(validation_alias="date") + + deployed: bool = Field(validation_alias="deployed") + + content: MetricVersionContent = Field(validation_alias="content") diff --git a/src/honeyhive/_generated/models/MetricVersionContent.py b/src/honeyhive/_generated/models/MetricVersionContent.py new file mode 100644 index 00000000..0353b3b5 --- /dev/null +++ b/src/honeyhive/_generated/models/MetricVersionContent.py @@ -0,0 +1,57 @@ +from typing import Any, Dict, List, Optional, Union + +from pydantic import BaseModel, Field + +from .MetricVersionContentFilters import MetricVersionContentFilters + +__all__ = ["MetricVersionContent"] + + +class MetricVersionContent(BaseModel): + """ + MetricVersionContent model + Versioned snapshot of a metric definition. + """ + + model_config = { + "populate_by_name": True, + "validate_assignment": True, + "extra": "allow", + "protected_namespaces": (), + } + + name: str = Field(validation_alias="name") + + type: str = Field(validation_alias="type") + + criteria: str = Field(validation_alias="criteria") + + description: Optional[str] = Field(validation_alias="description") + + return_type: str = Field(validation_alias="return_type") + + enabled_in_prod: bool = Field(validation_alias="enabled_in_prod") + + needs_ground_truth: bool = Field(validation_alias="needs_ground_truth") + + sampling_percentage: float = Field(validation_alias="sampling_percentage") + + model_provider: Optional[str] = Field( + validation_alias="model_provider", default=None + ) + + model_name: Optional[str] = Field(validation_alias="model_name", default=None) + + scale: Optional[int] = Field(validation_alias="scale", default=None) + + threshold: Optional[Dict[str, Any]] = Field( + validation_alias="threshold", default=None + ) + + categories: Optional[List[Any]] = Field(validation_alias="categories", default=None) + + child_metrics: Optional[List[Any]] = Field( + validation_alias="child_metrics", default=None + ) + + filters: MetricVersionContentFilters = Field(validation_alias="filters") diff --git a/src/honeyhive/_generated/models/MetricVersionContentCategoriesItem.py b/src/honeyhive/_generated/models/MetricVersionContentCategoriesItem.py new file mode 100644 index 00000000..3076b29c --- /dev/null +++ b/src/honeyhive/_generated/models/MetricVersionContentCategoriesItem.py @@ -0,0 +1,22 @@ +from typing import Any, Dict, List, Optional, Union + +from pydantic import BaseModel, Field + +__all__ = ["MetricVersionContentCategoriesItem"] + + +class MetricVersionContentCategoriesItem(BaseModel): + """ + MetricVersionContentCategoriesItem model + """ + + model_config = { + "populate_by_name": True, + "validate_assignment": True, + "extra": "allow", + "protected_namespaces": (), + } + + category: str = Field(validation_alias="category") + + score: Optional[float] = Field(validation_alias="score") diff --git a/src/honeyhive/_generated/models/MetricVersionContentChildMetricsItem.py b/src/honeyhive/_generated/models/MetricVersionContentChildMetricsItem.py new file mode 100644 index 00000000..ed0695e7 --- /dev/null +++ b/src/honeyhive/_generated/models/MetricVersionContentChildMetricsItem.py @@ -0,0 +1,26 @@ +from typing import Any, Dict, List, Optional, Union + +from pydantic import BaseModel, Field + +__all__ = ["MetricVersionContentChildMetricsItem"] + + +class MetricVersionContentChildMetricsItem(BaseModel): + """ + MetricVersionContentChildMetricsItem model + """ + + model_config = { + "populate_by_name": True, + "validate_assignment": True, + "extra": "allow", + "protected_namespaces": (), + } + + id: Optional[str] = Field(validation_alias="id", default=None) + + name: str = Field(validation_alias="name") + + weight: float = Field(validation_alias="weight") + + scale: Optional[int] = Field(validation_alias="scale", default=None) diff --git a/src/honeyhive/_generated/models/MetricVersionContentFilters.py b/src/honeyhive/_generated/models/MetricVersionContentFilters.py new file mode 100644 index 00000000..01baa947 --- /dev/null +++ b/src/honeyhive/_generated/models/MetricVersionContentFilters.py @@ -0,0 +1,22 @@ +from typing import Any, Dict, List, Optional, Union + +from pydantic import BaseModel, Field + +from .FiltersArray import FiltersArray + +__all__ = ["MetricVersionContentFilters"] + + +class MetricVersionContentFilters(BaseModel): + """ + MetricVersionContentFilters model + """ + + model_config = { + "populate_by_name": True, + "validate_assignment": True, + "extra": "allow", + "protected_namespaces": (), + } + + filterArray: FiltersArray = Field(validation_alias="filterArray") diff --git a/src/honeyhive/_generated/models/MetricVersionContentRequest.py b/src/honeyhive/_generated/models/MetricVersionContentRequest.py new file mode 100644 index 00000000..4dadfb54 --- /dev/null +++ b/src/honeyhive/_generated/models/MetricVersionContentRequest.py @@ -0,0 +1,72 @@ +from typing import Any, Dict, List, Optional, Union + +from pydantic import BaseModel, Field + +from .MetricVersionContentRequestFilters import MetricVersionContentRequestFilters + +__all__ = ["MetricVersionContentRequest"] + + +class MetricVersionContentRequest(BaseModel): + """ + MetricVersionContentRequest model + Metric definition snapshot accepted by POST /v1/metrics/{metric_id}/versions. + Six fields are optional and fall back to server-side defaults when omitted: + - `description` → `""` + - `return_type` → `"float"` + - `enabled_in_prod` → `true` for HUMAN metrics, `false` otherwise + - `needs_ground_truth` → `false` + - `sampling_percentage` → `10` + - `filters` → `{ "filterArray": [] }` + """ + + model_config = { + "populate_by_name": True, + "validate_assignment": True, + "extra": "allow", + "protected_namespaces": (), + } + + name: str = Field(validation_alias="name") + + type: str = Field(validation_alias="type") + + criteria: str = Field(validation_alias="criteria") + + description: Optional[str] = Field(validation_alias="description", default=None) + + return_type: Optional[str] = Field(validation_alias="return_type", default=None) + + enabled_in_prod: Optional[bool] = Field( + validation_alias="enabled_in_prod", default=None + ) + + needs_ground_truth: Optional[bool] = Field( + validation_alias="needs_ground_truth", default=None + ) + + sampling_percentage: Optional[float] = Field( + validation_alias="sampling_percentage", default=None + ) + + model_provider: Optional[str] = Field( + validation_alias="model_provider", default=None + ) + + model_name: Optional[str] = Field(validation_alias="model_name", default=None) + + scale: Optional[int] = Field(validation_alias="scale", default=None) + + threshold: Optional[Dict[str, Any]] = Field( + validation_alias="threshold", default=None + ) + + categories: Optional[List[Any]] = Field(validation_alias="categories", default=None) + + child_metrics: Optional[List[Any]] = Field( + validation_alias="child_metrics", default=None + ) + + filters: Optional[MetricVersionContentRequestFilters] = Field( + validation_alias="filters", default=None + ) diff --git a/src/honeyhive/_generated/models/MetricVersionContentRequestCategoriesItem.py b/src/honeyhive/_generated/models/MetricVersionContentRequestCategoriesItem.py new file mode 100644 index 00000000..c608f67a --- /dev/null +++ b/src/honeyhive/_generated/models/MetricVersionContentRequestCategoriesItem.py @@ -0,0 +1,22 @@ +from typing import Any, Dict, List, Optional, Union + +from pydantic import BaseModel, Field + +__all__ = ["MetricVersionContentRequestCategoriesItem"] + + +class MetricVersionContentRequestCategoriesItem(BaseModel): + """ + MetricVersionContentRequestCategoriesItem model + """ + + model_config = { + "populate_by_name": True, + "validate_assignment": True, + "extra": "allow", + "protected_namespaces": (), + } + + category: str = Field(validation_alias="category") + + score: Optional[float] = Field(validation_alias="score") diff --git a/src/honeyhive/_generated/models/MetricVersionContentRequestChildMetricsItem.py b/src/honeyhive/_generated/models/MetricVersionContentRequestChildMetricsItem.py new file mode 100644 index 00000000..1f3837c3 --- /dev/null +++ b/src/honeyhive/_generated/models/MetricVersionContentRequestChildMetricsItem.py @@ -0,0 +1,26 @@ +from typing import Any, Dict, List, Optional, Union + +from pydantic import BaseModel, Field + +__all__ = ["MetricVersionContentRequestChildMetricsItem"] + + +class MetricVersionContentRequestChildMetricsItem(BaseModel): + """ + MetricVersionContentRequestChildMetricsItem model + """ + + model_config = { + "populate_by_name": True, + "validate_assignment": True, + "extra": "allow", + "protected_namespaces": (), + } + + id: Optional[str] = Field(validation_alias="id", default=None) + + name: str = Field(validation_alias="name") + + weight: float = Field(validation_alias="weight") + + scale: Optional[int] = Field(validation_alias="scale", default=None) diff --git a/src/honeyhive/_generated/models/MetricVersionContentRequestFilters.py b/src/honeyhive/_generated/models/MetricVersionContentRequestFilters.py new file mode 100644 index 00000000..bf5f2a02 --- /dev/null +++ b/src/honeyhive/_generated/models/MetricVersionContentRequestFilters.py @@ -0,0 +1,23 @@ +from typing import Any, Dict, List, Optional, Union + +from pydantic import BaseModel, Field + +from .FiltersArray import FiltersArray + +__all__ = ["MetricVersionContentRequestFilters"] + + +class MetricVersionContentRequestFilters(BaseModel): + """ + MetricVersionContentRequestFilters model + ETL filter narrowing which events this metric applies to. Defaults to `{ filterArray: [] }` (no filtering). + """ + + model_config = { + "populate_by_name": True, + "validate_assignment": True, + "extra": "allow", + "protected_namespaces": (), + } + + filterArray: FiltersArray = Field(validation_alias="filterArray") diff --git a/src/honeyhive/_generated/models/__init__.py b/src/honeyhive/_generated/models/__init__.py index a2d408f8..9322d58a 100644 --- a/src/honeyhive/_generated/models/__init__.py +++ b/src/honeyhive/_generated/models/__init__.py @@ -33,6 +33,9 @@ from .CreateMetricRequestChildMetricsItem import * from .CreateMetricRequestFilters import * from .CreateMetricResponse import * +from .CreateMetricVersionParams import * +from .CreateMetricVersionRequest import * +from .CreateMetricVersionResponse import * from .Datapoint import * from .DatapointMapping import * from .DatapointResult import * @@ -49,6 +52,8 @@ from .DeleteMetricParams import * from .DeleteMetricResponse import * from .DeleteResult import * +from .DeployMetricVersionParams import * +from .DeployMetricVersionResponse import * from .Event import * from .EventComparisonDetail import * from .EventDetail import * @@ -99,6 +104,8 @@ from .GetExperimentRunSummaryQuery import * from .GetMetricsQuery import * from .GetMetricsResponse import * +from .GetMetricVersionsParams import * +from .GetMetricVersionsResponse import * from .GetProjectsResponse import * from .GetRunSchemaDateRangeOneOf1 import * from .GetRunSchemaQuery import * @@ -143,6 +150,15 @@ from .MetricItemChildMetricsItem import * from .MetricItemFilters import * from .MetricsAggregation import * +from .MetricVersion import * +from .MetricVersionContent import * +from .MetricVersionContentCategoriesItem import * +from .MetricVersionContentChildMetricsItem import * +from .MetricVersionContentFilters import * +from .MetricVersionContentRequest import * +from .MetricVersionContentRequestCategoriesItem import * +from .MetricVersionContentRequestChildMetricsItem import * +from .MetricVersionContentRequestFilters import * from .ModelEvent import * from .Pagination import * from .PassingRange import * diff --git a/src/honeyhive/_generated/services/Metric_Versions_service.py b/src/honeyhive/_generated/services/Metric_Versions_service.py new file mode 100644 index 00000000..e3307da4 --- /dev/null +++ b/src/honeyhive/_generated/services/Metric_Versions_service.py @@ -0,0 +1,124 @@ +from typing import * + +import httpx + +from ..api_config import APIConfig, HTTPException, _serialize_query_params +from ..models import * + + +def getMetricVersions( + api_config_override: Optional[APIConfig] = None, *, metric_id: str +) -> GetMetricVersionsResponse: + api_config = api_config_override if api_config_override else APIConfig() + + base_path = api_config.base_path + path = f"/v1/metrics/{metric_id}/versions" + headers = api_config.get_default_headers() + query_params: Dict[str, Any] = {} + + query_params = { + key: value for (key, value) in query_params.items() if value is not None + } + + with httpx.Client(base_url=base_path, verify=api_config.verify) as client: + response = client.request( + "get", + httpx.URL(path), + headers=headers, + params=_serialize_query_params(query_params), + ) + + if response.status_code != 200: + raise HTTPException( + response.status_code, + f"getMetricVersions failed with status code: {response.status_code}", + ) + else: + body = None if 200 == 204 else response.json() + + return ( + GetMetricVersionsResponse(**body) + if body is not None + else GetMetricVersionsResponse() + ) + + +def createMetricVersion( + api_config_override: Optional[APIConfig] = None, + *, + metric_id: str, + data: CreateMetricVersionRequest, +) -> CreateMetricVersionResponse: + api_config = api_config_override if api_config_override else APIConfig() + + base_path = api_config.base_path + path = f"/v1/metrics/{metric_id}/versions" + headers = api_config.get_default_headers() + query_params: Dict[str, Any] = {} + + query_params = { + key: value for (key, value) in query_params.items() if value is not None + } + + with httpx.Client(base_url=base_path, verify=api_config.verify) as client: + response = client.request( + "post", + httpx.URL(path), + headers=headers, + params=_serialize_query_params(query_params), + json=data.model_dump(exclude_none=True), + ) + + if response.status_code != 200: + raise HTTPException( + response.status_code, + f"createMetricVersion failed with status code: {response.status_code}", + ) + else: + body = None if 200 == 204 else response.json() + + return ( + CreateMetricVersionResponse(**body) + if body is not None + else CreateMetricVersionResponse() + ) + + +def deployMetricVersion( + api_config_override: Optional[APIConfig] = None, + *, + metric_id: str, + version_name: str, +) -> DeployMetricVersionResponse: + api_config = api_config_override if api_config_override else APIConfig() + + base_path = api_config.base_path + path = f"/v1/metrics/{metric_id}/versions/{version_name}/deploy" + headers = api_config.get_default_headers() + query_params: Dict[str, Any] = {} + + query_params = { + key: value for (key, value) in query_params.items() if value is not None + } + + with httpx.Client(base_url=base_path, verify=api_config.verify) as client: + response = client.request( + "post", + httpx.URL(path), + headers=headers, + params=_serialize_query_params(query_params), + ) + + if response.status_code != 200: + raise HTTPException( + response.status_code, + f"deployMetricVersion failed with status code: {response.status_code}", + ) + else: + body = None if 200 == 204 else response.json() + + return ( + DeployMetricVersionResponse(**body) + if body is not None + else DeployMetricVersionResponse() + ) diff --git a/src/honeyhive/_generated/services/async_Metric_Versions_service.py b/src/honeyhive/_generated/services/async_Metric_Versions_service.py new file mode 100644 index 00000000..718b99c0 --- /dev/null +++ b/src/honeyhive/_generated/services/async_Metric_Versions_service.py @@ -0,0 +1,130 @@ +from typing import * + +import httpx + +from ..api_config import APIConfig, HTTPException, _serialize_query_params +from ..models import * + + +async def getMetricVersions( + api_config_override: Optional[APIConfig] = None, *, metric_id: str +) -> GetMetricVersionsResponse: + api_config = api_config_override if api_config_override else APIConfig() + + base_path = api_config.base_path + path = f"/v1/metrics/{metric_id}/versions" + headers = api_config.get_default_headers() + query_params: Dict[str, Any] = {} + + query_params = { + key: value for (key, value) in query_params.items() if value is not None + } + + async with httpx.AsyncClient( + base_url=base_path, verify=api_config.verify + ) as client: + response = await client.request( + "get", + httpx.URL(path), + headers=headers, + params=_serialize_query_params(query_params), + ) + + if response.status_code != 200: + raise HTTPException( + response.status_code, + f"getMetricVersions failed with status code: {response.status_code}", + ) + else: + body = None if 200 == 204 else response.json() + + return ( + GetMetricVersionsResponse(**body) + if body is not None + else GetMetricVersionsResponse() + ) + + +async def createMetricVersion( + api_config_override: Optional[APIConfig] = None, + *, + metric_id: str, + data: CreateMetricVersionRequest, +) -> CreateMetricVersionResponse: + api_config = api_config_override if api_config_override else APIConfig() + + base_path = api_config.base_path + path = f"/v1/metrics/{metric_id}/versions" + headers = api_config.get_default_headers() + query_params: Dict[str, Any] = {} + + query_params = { + key: value for (key, value) in query_params.items() if value is not None + } + + async with httpx.AsyncClient( + base_url=base_path, verify=api_config.verify + ) as client: + response = await client.request( + "post", + httpx.URL(path), + headers=headers, + params=_serialize_query_params(query_params), + json=data.model_dump(exclude_none=True), + ) + + if response.status_code != 200: + raise HTTPException( + response.status_code, + f"createMetricVersion failed with status code: {response.status_code}", + ) + else: + body = None if 200 == 204 else response.json() + + return ( + CreateMetricVersionResponse(**body) + if body is not None + else CreateMetricVersionResponse() + ) + + +async def deployMetricVersion( + api_config_override: Optional[APIConfig] = None, + *, + metric_id: str, + version_name: str, +) -> DeployMetricVersionResponse: + api_config = api_config_override if api_config_override else APIConfig() + + base_path = api_config.base_path + path = f"/v1/metrics/{metric_id}/versions/{version_name}/deploy" + headers = api_config.get_default_headers() + query_params: Dict[str, Any] = {} + + query_params = { + key: value for (key, value) in query_params.items() if value is not None + } + + async with httpx.AsyncClient( + base_url=base_path, verify=api_config.verify + ) as client: + response = await client.request( + "post", + httpx.URL(path), + headers=headers, + params=_serialize_query_params(query_params), + ) + + if response.status_code != 200: + raise HTTPException( + response.status_code, + f"deployMetricVersion failed with status code: {response.status_code}", + ) + else: + body = None if 200 == 204 else response.json() + + return ( + DeployMetricVersionResponse(**body) + if body is not None + else DeployMetricVersionResponse() + ) diff --git a/src/honeyhive/api/__init__.py b/src/honeyhive/api/__init__.py index 541bf1dd..9638f2f6 100644 --- a/src/honeyhive/api/__init__.py +++ b/src/honeyhive/api/__init__.py @@ -15,6 +15,7 @@ ExperimentsAPI, HoneyHive, MetricsAPI, + MetricVersionsAPI, SessionsAPI, ) @@ -31,6 +32,7 @@ "EventsAPI", "ExperimentsAPI", "MetricsAPI", + "MetricVersionsAPI", "SessionsAPI", # Backwards compatible aliases "EvaluationsAPI", diff --git a/src/honeyhive/api/client.py b/src/honeyhive/api/client.py index a350095d..f6dcd585 100644 --- a/src/honeyhive/api/client.py +++ b/src/honeyhive/api/client.py @@ -39,6 +39,7 @@ from honeyhive._generated.services import Datasets_service as datasets_svc from honeyhive._generated.services import Events_service as events_svc from honeyhive._generated.services import Experiments_service as experiments_svc +from honeyhive._generated.services import Metric_Versions_service as metric_versions_svc from honeyhive._generated.services import Metrics_service as metrics_svc from honeyhive._generated.services import Sessions_service as sessions_svc from honeyhive._generated.services import async_Charts_service as charts_svc_async @@ -53,6 +54,9 @@ from honeyhive._generated.services import ( async_Experiments_service as experiments_svc_async, ) +from honeyhive._generated.services import ( + async_Metric_Versions_service as metric_versions_svc_async, +) from honeyhive._generated.services import async_Metrics_service as metrics_svc_async from honeyhive._generated.services import async_Sessions_service as sessions_svc_async @@ -71,12 +75,15 @@ CreateDatasetResponse, CreateMetricRequest, CreateMetricResponse, + CreateMetricVersionRequest, + CreateMetricVersionResponse, DeleteChartResponse, DeleteConfigurationResponse, DeleteDatapointResponse, DeleteDatasetResponse, DeleteExperimentRunResponse, DeleteMetricResponse, + DeployMetricVersionResponse, EventExportResponse, EventFilter, GetChartResponse, @@ -89,6 +96,7 @@ GetEventsSchemaResponse, GetExperimentRunResponse, GetExperimentRunsResponse, + GetMetricVersionsResponse, MetricItem, Pagination, PostEventBatchRequest, @@ -1881,6 +1889,63 @@ def list_metrics( return self.list(project=project, name=name, type=type) +class MetricVersionsAPI(BaseAPI): + """Metric Versions API. + + Each metric has an immutable history of versions; one version at a time is + marked as deployed (used by evaluations). Versions are nested under the + parent metric at ``/v1/metrics/{metric_id}/versions``. + """ + + # Sync methods + def list(self, metric_id: str) -> GetMetricVersionsResponse: + """List all versions for a metric, oldest-first.""" + return metric_versions_svc.getMetricVersions( + self._api_config, metric_id=metric_id + ) + + def create( + self, metric_id: str, request: CreateMetricVersionRequest + ) -> CreateMetricVersionResponse: + """Create a new version of a metric. + + Set ``request.deploy_immediately = True`` to atomically mark the new + version as deployed in the same transaction. + """ + return metric_versions_svc.createMetricVersion( + self._api_config, metric_id=metric_id, data=request + ) + + def deploy(self, metric_id: str, version_name: str) -> DeployMetricVersionResponse: + """Deploy an existing version of a metric, replacing the live version.""" + return metric_versions_svc.deployMetricVersion( + self._api_config, metric_id=metric_id, version_name=version_name + ) + + # Async methods + async def list_async(self, metric_id: str) -> GetMetricVersionsResponse: + """List all versions for a metric asynchronously.""" + return await metric_versions_svc_async.getMetricVersions( + self._api_config, metric_id=metric_id + ) + + async def create_async( + self, metric_id: str, request: CreateMetricVersionRequest + ) -> CreateMetricVersionResponse: + """Create a new version of a metric asynchronously.""" + return await metric_versions_svc_async.createMetricVersion( + self._api_config, metric_id=metric_id, data=request + ) + + async def deploy_async( + self, metric_id: str, version_name: str + ) -> DeployMetricVersionResponse: + """Deploy an existing version of a metric asynchronously.""" + return await metric_versions_svc_async.deployMetricVersion( + self._api_config, metric_id=metric_id, version_name=version_name + ) + + class SessionsAPI(BaseAPI): """Sessions API.""" @@ -2064,6 +2129,7 @@ def __init__( self.events = EventsAPI(self._api_config) self.experiments = ExperimentsAPI(self._api_config) self.metrics = MetricsAPI(self._api_config) + self.metric_versions = MetricVersionsAPI(self._api_config) self.sessions = SessionsAPI(self._api_config) # Alias for backwards compatibility diff --git a/src/honeyhive/models/__init__.py b/src/honeyhive/models/__init__.py index e262484e..d0c83f66 100644 --- a/src/honeyhive/models/__init__.py +++ b/src/honeyhive/models/__init__.py @@ -217,6 +217,8 @@ class EventExportResponse(BaseModel): CreateDatasetResponse, CreateMetricRequest, CreateMetricResponse, + CreateMetricVersionRequest, + CreateMetricVersionResponse, DatapointMapping, DeleteChartResponse, DeleteConfigurationResponse, @@ -228,6 +230,7 @@ class EventExportResponse(BaseModel): DeleteExperimentRunParams, DeleteExperimentRunResponse, DeleteMetricResponse, + DeployMetricVersionResponse, Event, GetChartResponse, GetChartsResponse, @@ -255,7 +258,11 @@ class EventExportResponse(BaseModel): GetExperimentRunsSchemaResponse, GetMetricsQuery, GetMetricsResponse, + GetMetricVersionsResponse, MetricItem, + MetricVersion, + MetricVersionContent, + MetricVersionContentRequest, Pagination, PostEventBatchRequest, PostEventBatchResponse, @@ -374,6 +381,14 @@ class EventExportResponse(BaseModel): "RunMetricResponse", "UpdateMetricRequest", "UpdateMetricResponse", + # Metric version models + "CreateMetricVersionRequest", + "CreateMetricVersionResponse", + "DeployMetricVersionResponse", + "GetMetricVersionsResponse", + "MetricVersion", + "MetricVersionContent", + "MetricVersionContentRequest", # Session models "PostSessionRequest", "PostSessionStartResponse", diff --git a/src/honeyhive/models/models.py b/src/honeyhive/models/models.py index 2da82d91..0189397d 100644 --- a/src/honeyhive/models/models.py +++ b/src/honeyhive/models/models.py @@ -26,6 +26,8 @@ CreateDatasetResponse, CreateMetricRequest, CreateMetricResponse, + CreateMetricVersionRequest, + CreateMetricVersionResponse, DatapointMapping, DeleteChartResponse, DeleteConfigurationResponse, @@ -36,6 +38,7 @@ DeleteExperimentRunParams, DeleteExperimentRunResponse, DeleteMetricResponse, + DeployMetricVersionResponse, Event, GetChartResponse, GetChartsResponse, @@ -56,6 +59,7 @@ GetExperimentRunsResponse, GetMetricsQuery, GetMetricsResponse, + GetMetricVersionsResponse, LegacyDeleteDatasetQuery, LegacyEvent, LegacyGetEventsSchemaQuery, @@ -71,6 +75,9 @@ LegacyUpdateEventRequest, LegacyUpdateMetricRequest, MetricItem, + MetricVersion, + MetricVersionContent, + MetricVersionContentRequest, Pagination, PostEventBatchRequest, PostEventBatchResponse, @@ -246,6 +253,8 @@ class UpdateEventRequest(LegacyUpdateEventRequest): "CreateDatasetResponse", "CreateMetricRequest", "CreateMetricResponse", + "CreateMetricVersionRequest", + "CreateMetricVersionResponse", "DatapointMapping", "DeleteChartResponse", "DeleteConfigurationResponse", @@ -257,6 +266,7 @@ class UpdateEventRequest(LegacyUpdateEventRequest): "DeleteExperimentRunParams", "DeleteExperimentRunResponse", "DeleteMetricResponse", + "DeployMetricVersionResponse", "Event", "GetChartResponse", "GetChartsResponse", @@ -284,8 +294,12 @@ class UpdateEventRequest(LegacyUpdateEventRequest): "GetExperimentRunsSchemaResponse", "GetMetricsQuery", "GetMetricsResponse", + "GetMetricVersionsResponse", "LegacyEvent", "MetricItem", + "MetricVersion", + "MetricVersionContent", + "MetricVersionContentRequest", "Pagination", "PostEventBatchRequest", "PostEventBatchResponse", diff --git a/src/honeyhive/tracer/instrumentation/decorators.py b/src/honeyhive/tracer/instrumentation/decorators.py index 69442c44..c188868c 100644 --- a/src/honeyhive/tracer/instrumentation/decorators.py +++ b/src/honeyhive/tracer/instrumentation/decorators.py @@ -213,16 +213,14 @@ def _capture_function_inputs( span.set_attribute(f"honeyhive_inputs.{param_name}", param_value) elif isinstance(param_value, (dict, list)): # Complex types: JSON serialize + # TODO: support threading through JSON structures here so we + # can serialize to OTLP kvlist_value: + # https://opentelemetry.io/docs/specs/otel/common/attribute-type-mapping/#associative-arrays-with-unique-keys serialized = json.dumps(param_value) - # Truncate if too long (prevent huge spans) - if len(serialized) > 1000: - serialized = serialized[:1000] + "... (truncated)" span.set_attribute(f"honeyhive_inputs.{param_name}", serialized) else: # Other types: use str() representation str_value = str(param_value) - if len(str_value) > 500: - str_value = str_value[:500] + "... (truncated)" span.set_attribute(f"honeyhive_inputs.{param_name}", str_value) except Exception: # Skip non-serializable values silently diff --git a/src/honeyhive/tracer/processing/otlp_exporter.py b/src/honeyhive/tracer/processing/otlp_exporter.py index 3fb4fe82..501aec1c 100644 --- a/src/honeyhive/tracer/processing/otlp_exporter.py +++ b/src/honeyhive/tracer/processing/otlp_exporter.py @@ -98,7 +98,10 @@ def _to_otlp_any_value(cls, value: Any) -> Dict[str, Any]: if isinstance(value, bool): return {"boolValue": value} if isinstance(value, int): - return {"intValue": value} + # protobuf JSON mapping: int64 must be a JSON string so values above + # 2^53 survive the server's float64 decode path without precision loss. + # Matches native opentelemetry-exporter-otlp-proto-http behavior. + return {"intValue": str(value)} if isinstance(value, float): # NaN/Inf are not valid JSON — Python's json.dumps emits them as # `NaN`/`Infinity` tokens, which Go's encoding/json rejects and @@ -113,6 +116,9 @@ def _to_otlp_any_value(cls, value: Any) -> Dict[str, Any]: return { "arrayValue": {"values": [cls._to_otlp_any_value(v) for v in value]} } + # TODO: serialize JSON to kvlist_value here so it doesnt fall through + # to stringValue: + # https://opentelemetry.io/docs/specs/otel/common/attribute-type-mapping/#associative-arrays-with-unique-keys return {"stringValue": str(value)} @classmethod @@ -154,8 +160,8 @@ def _span_to_otlp_json(self, span: ReadableSpan) -> Dict[str, Any]: event_attrs = self._to_otlp_key_values(event.attributes) events.append( { - # uint64 - nanoseconds since Unix epoch - "timeUnixNano": event.timestamp, + # uint64 - protobuf JSON mapping requires string for uint64 + "timeUnixNano": str(event.timestamp), "name": event.name, "attributes": event_attrs, } @@ -185,9 +191,9 @@ def _span_to_otlp_json(self, span: ReadableSpan) -> Dict[str, Any]: "parentSpanId": parent_span_id, "name": span.name, "kind": span_kind, - # uint64 - nanoseconds since Unix epoch - "startTimeUnixNano": span.start_time, - "endTimeUnixNano": span.end_time, + # uint64 - protobuf JSON mapping requires string for uint64 + "startTimeUnixNano": str(span.start_time), + "endTimeUnixNano": str(span.end_time), "attributes": attributes, "events": events, "status": status, diff --git a/tests/integration/_mock_llm_helpers.py b/tests/integration/_mock_llm_helpers.py new file mode 100644 index 00000000..acfb60a9 --- /dev/null +++ b/tests/integration/_mock_llm_helpers.py @@ -0,0 +1,50 @@ +"""Shared helpers for tests that drive the mock-llm service through the +OpenAI Python SDK. + +The leading underscore keeps pytest from collecting this file as a test +module. Imported by integration tests that swap out real OpenAI for +mock-llm's OpenAI-compatible endpoint at ``http://localhost:9025/v1``. +This mirrors the centralization pattern already used by +``_experiments_helpers.require_server_side_eval_creds`` so all mock-llm +config lives in one place. +""" + +from __future__ import annotations + +import os +from typing import TYPE_CHECKING + +if TYPE_CHECKING: + from openai import OpenAI + +# Path-style model targets mock-llm's ``deterministic`` provider, which has +# no latency and ``error_rate: 0``. See mock_provider_config.yaml. +MOCK_LLM_MODEL = "deterministic/default/default" +MOCK_LLM_BASE_URL = "http://localhost:9025/v1" + +# Mirror of ``mock_llm_service.main.MOCK_RESPONSE``. Kept here so tests can +# assert exact equality against the mocked response without reaching into the +# service package. If the service's constant changes, the streaming test will +# fail loudly — that's intentional: it surfaces the drift rather than letting +# it silently weaken the assertion. +MOCK_RESPONSE = "This is a mock explanation. Rating: [[5]]" + + +def mock_openai_client() -> OpenAI: + """Return an OpenAI SDK client wired to mock-llm. + + Reads ``MOCK_LLM_API_KEY`` from env with a fallback to the literal + that mock-llm's container starts with in ``config/dev.env``, matching + the pattern used by the TS mock-llm tests. + + The ``openai`` import is deferred to the call site so importing this + helper does not promote ``openai`` to a collection-time hard dep — + callers can still use ``pytest.importorskip("openai")`` at the test + level for environments without the ``openinference-openai`` extra. + """ + from openai import OpenAI + + return OpenAI( + api_key=os.environ.get("MOCK_LLM_API_KEY", "mock-llm-api-key"), + base_url=MOCK_LLM_BASE_URL, + ) diff --git a/tests/integration/api/test_tools_api.py b/tests/integration/api/test_tools_api.py deleted file mode 100644 index 1db17392..00000000 --- a/tests/integration/api/test_tools_api.py +++ /dev/null @@ -1,308 +0,0 @@ -"""ToolsAPI Integration Tests - NO MOCKS, REAL API CALLS. - -NOTE: Tool models (CreateToolRequest, CreateToolResponse, etc.) were removed -from the generated SDK models. These tests are skipped until the OpenAPI spec -is updated and the models are regenerated. -""" - -import time -import uuid -from typing import Any - -import pytest - -# Tool models have been removed from the generated SDK. -# Guard the import so the file can be collected without ImportError. -try: - from honeyhive.models import ( - CreateToolRequest, - CreateToolResponse, - DeleteToolResponse, - GetToolsResponse, - UpdateToolRequest, - UpdateToolResponse, - ) - - _TOOL_MODELS_AVAILABLE = True -except ImportError: - _TOOL_MODELS_AVAILABLE = False - -pytestmark = pytest.mark.skipif( - not _TOOL_MODELS_AVAILABLE, - reason="Tool models removed from SDK — waiting for OpenAPI spec update", -) - - -class TestToolsAPI: - """Test ToolsAPI CRUD operations. - - Note: Several tests are skipped due to discovered client-level bugs: - - tools.delete() has a bug where the client wrapper passes 'tool_id=id' but the - generated service expects 'function_id' parameter. This is a client wrapper bug. - - tools.update() returns 400 error from the backend. - These issues should be fixed in the client wrapper. - """ - - def test_create_tool( - self, integration_client: Any, integration_project_name: str - ) -> None: - """Test tool creation with schema and parameters, verify backend storage.""" - test_id = str(uuid.uuid4())[:8] - tool_name = f"test_tool_{test_id}" - - tool_request = CreateToolRequest( - name=tool_name, - description=f"Integration test tool {test_id}", - parameters={ - "type": "function", - "function": { - "name": tool_name, - "description": "Test function", - "parameters": { - "type": "object", - "properties": { - "query": {"type": "string", "description": "Search query"} - }, - "required": ["query"], - }, - }, - }, - tool_type="function", - ) - - response = integration_client.tools.create(tool_request) - - # Verify response is CreateToolResponse with inserted and result fields - assert isinstance(response, CreateToolResponse) - assert response.inserted is True - # Tools API returns id directly in result, not insertedIds - assert "id" in response.result - tool_id = response.result["id"] - assert tool_id is not None - - # Note: Cleanup removed - tools.delete() has a bug where client wrapper - # passes 'tool_id' but generated service expects 'function_id' parameter - - @pytest.mark.skip( - reason="Client Bug: tools.delete() passes tool_id but service expects function_id - cleanup would fail" - ) - def test_get_tool( - self, integration_client: Any, integration_project_name: str - ) -> None: - """Test retrieval by ID, verify schema intact.""" - test_id = str(uuid.uuid4())[:8] - tool_name = f"test_get_tool_{test_id}" - - # Create a tool first - tool_request = CreateToolRequest( - name=tool_name, - description=f"Integration test tool for retrieval {test_id}", - parameters={ - "type": "function", - "function": { - "name": tool_name, - "description": "Test function", - "parameters": { - "type": "object", - "properties": { - "query": {"type": "string", "description": "Search query"} - }, - "required": ["query"], - }, - }, - }, - tool_type="function", - ) - - create_resp = integration_client.tools.create(tool_request) - assert isinstance(create_resp, CreateToolResponse) - assert create_resp.inserted is True - # Tools API returns id directly in result - assert "id" in create_resp.result - tool_id = create_resp.result["id"] - - # Wait for indexing - time.sleep(2) - - # v1 API doesn't have a direct get method, use list and filter - tools_list = integration_client.tools.list() - assert isinstance(tools_list, list) - - # Find the created tool by ID - retrieved_tool = None - for tool in tools_list: - # GetToolsResponse is a dynamic Pydantic model, access fields via model_dump() - tool_dict = tool.model_dump() - # Check for id or _id field (backend may use either) - tool_id_from_response = tool_dict.get("id") or tool_dict.get("_id") - if tool_id_from_response == tool_id: - retrieved_tool = tool_dict - break - - assert retrieved_tool is not None - assert retrieved_tool.get("name") == tool_name - - # Note: Cleanup removed - tools.delete() has a bug where client wrapper - # passes 'tool_id' but generated service expects 'function_id' parameter - - def test_get_tool_404(self, integration_client: Any) -> None: - """Test 404 for missing tool (v1 API doesn't have get_tool method).""" - pytest.skip("v1 API doesn't have get_tool method, only list") - - @pytest.mark.skip( - reason="Client Bug: tools.delete() passes tool_id but service expects function_id - cleanup would fail" - ) - def test_list_tools( - self, integration_client: Any, integration_project_name: str - ) -> None: - """Test listing with project filtering, pagination.""" - test_id = str(uuid.uuid4())[:8] - tool_ids = [] - - # Create 2-3 tools - for i in range(3): - tool_name = f"test_list_tool_{test_id}_{i}" - tool_request = CreateToolRequest( - name=tool_name, - description=f"Integration test tool {i} for listing {test_id}", - parameters={ - "type": "function", - "function": { - "name": tool_name, - "description": f"Test function {i}", - "parameters": { - "type": "object", - "properties": { - "query": { - "type": "string", - "description": "Search query", - } - }, - "required": ["query"], - }, - }, - }, - tool_type="function", - ) - - create_resp = integration_client.tools.create(tool_request) - assert isinstance(create_resp, CreateToolResponse) - assert create_resp.inserted is True - # Tools API returns id directly in result - assert "id" in create_resp.result - tool_ids.append(create_resp.result["id"]) - - # Wait for indexing - time.sleep(2) - - # Call client.tools.list() - tools_list = integration_client.tools.list() - - # Verify we get a list response - assert isinstance(tools_list, list) - # May be empty or contain tools, that's ok - basic existence check - assert len(tools_list) >= 0 - - # Note: Cleanup removed - tools.delete() has a bug where client wrapper - # passes 'tool_id' but generated service expects 'function_id' parameter - - @pytest.mark.skip(reason="Backend returns 400 error for updateTool endpoint") - def test_update_tool( - self, integration_client: Any, integration_project_name: str - ) -> None: - """Test tool schema updates, parameter changes, verify persistence.""" - test_id = str(uuid.uuid4())[:8] - tool_name = f"test_update_tool_{test_id}" - - # Create a tool - tool_request = CreateToolRequest( - name=tool_name, - description=f"Integration test tool {test_id}", - parameters={ - "type": "function", - "function": { - "name": tool_name, - "description": "Test function", - "parameters": { - "type": "object", - "properties": { - "query": {"type": "string", "description": "Search query"} - }, - "required": ["query"], - }, - }, - }, - tool_type="function", - ) - - create_resp = integration_client.tools.create(tool_request) - assert isinstance(create_resp, CreateToolResponse) - assert create_resp.inserted is True - # Tools API returns id directly in result - assert "id" in create_resp.result - tool_id = create_resp.result["id"] - - # Wait for indexing - time.sleep(2) - - # Create UpdateToolRequest with updated description - updated_description = f"Updated description {test_id}" - update_request = UpdateToolRequest(id=tool_id, description=updated_description) - - # Call client.tools.update(tool_id, update_request) - response = integration_client.tools.update(update_request) - - # Verify response - assert isinstance(response, UpdateToolResponse) - assert response.updated is True - - # Note: Cleanup removed - tools.delete() has a bug where client wrapper - # passes 'tool_id' but generated service expects 'function_id' parameter - - @pytest.mark.skip( - reason="Client Bug: tools.delete() passes tool_id but generated service expects function_id parameter" - ) - def test_delete_tool( - self, integration_client: Any, integration_project_name: str - ) -> None: - """Test deletion, verify not in list after delete.""" - test_id = str(uuid.uuid4())[:8] - tool_name = f"test_delete_tool_{test_id}" - - # Create a tool - tool_request = CreateToolRequest( - name=tool_name, - description=f"Integration test tool {test_id}", - parameters={ - "type": "function", - "function": { - "name": tool_name, - "description": "Test function", - "parameters": { - "type": "object", - "properties": { - "query": {"type": "string", "description": "Search query"} - }, - "required": ["query"], - }, - }, - }, - tool_type="function", - ) - - create_resp = integration_client.tools.create(tool_request) - assert isinstance(create_resp, CreateToolResponse) - assert create_resp.inserted is True - # Tools API returns id directly in result - assert "id" in create_resp.result - tool_id = create_resp.result["id"] - - # Wait for indexing - time.sleep(2) - - # Call client.tools.delete(tool_id) - response = integration_client.tools.delete(tool_id) - - # Verify response indicates deletion - assert isinstance(response, DeleteToolResponse) - assert response.deleted is True diff --git a/tests/integration/ci_known_failures.txt b/tests/integration/ci_known_failures.txt index 1ebaf728..bde61d79 100644 --- a/tests/integration/ci_known_failures.txt +++ b/tests/integration/ci_known_failures.txt @@ -12,22 +12,12 @@ # are now purely navigational — the real source of truth is per-entry. # # Remaining buckets: -# * OTLPJSONExporter int64 precision (HHAI-5004) — raw JSON number -# overflows float64 for |v| > 2^53, backend rejects the batch # * enrich_span attribute loss on error-raising spans — attrs set via # `enrich_span({...})` inside a function that raises never make it # onto the exported span (test.unique_id lookup fails) # * Session-id override routing — honeyhive.session_id attribute # override routes events to a different session than the one the # test queries under -# * Experiment comparison / semantic assertions - -# OTLPJSONExporter int64 precision bug (HHAI-5004) — the test sends -# attr.large_int=9223372036854775807; the exporter now emits it as a raw -# JSON number, but the Go backend decodes through float64, rounds up to -# 9223372036854776000, and rejects the payload with HTTP 500 (int64 -# overflow). Blocks the entire batch, so the event never ingests. -tests/integration/test_otel_backend_verification_integration.py::TestOTELBackendVerificationIntegration::test_high_cardinality_attributes_backend_verification # enrich_span attribute loss on error-raising spans — OTLP export # succeeds with HTTP 200, but attributes set via `enrich_span({...})` @@ -44,13 +34,3 @@ tests/integration/test_otel_backend_verification_integration.py::TestOTELBackend # stringification issue. tests/integration/test_otel_otlp_export_integration.py::TestOTELOTLPExportIntegration::test_otlp_export_with_backend_verification -# _enrich_session_with_results race against OTLP→writer pipeline — the -# session event is exported via OTLP at the end of evaluate(), then -# _enrich_session_with_results immediately PUTs ground_truth feedback -# onto it via events.update(). On the local CI stack the writer hasn't -# materialized the session row yet, so the PUT 404s -# ("updateEventLegacy failed with status code: 404") and feedback stays -# empty, tripping the TASK 3 assertion. Passes against testing-dp-1 -# where the writer keeps up. Fix is SDK-side: retry the enrich PUT with -# a short backoff, separate change. -tests/integration/test_v1_immediate_ship_requirements.py::TestV1ImmediateShipRequirements::test_all_five_requirements_end_to_end diff --git a/tests/integration/test_e2e_patterns.py b/tests/integration/test_e2e_patterns.py index 541a9bf2..07b3d0ef 100644 --- a/tests/integration/test_e2e_patterns.py +++ b/tests/integration/test_e2e_patterns.py @@ -18,6 +18,7 @@ from honeyhive import HoneyHiveTracer, enrich_span, trace from honeyhive.tracer.registry import set_default_tracer +from tests.integration._mock_llm_helpers import MOCK_LLM_MODEL, mock_openai_client # Skip all tests if no API key pytestmark = pytest.mark.skipif( @@ -158,11 +159,6 @@ def test_openai_with_enrichment(self) -> None: """Test OpenAI call with span enrichment.""" pytest.importorskip("openai") - from openai import ( # pylint: disable=import-outside-toplevel,import-error - OpenAI, - ) - - # Optional dependency with skip marker # Initialize tracer tracer = HoneyHiveTracer.init( api_key=os.environ["HH_API_KEY"], @@ -170,25 +166,25 @@ def test_openai_with_enrichment(self) -> None: session_name="e2e-test-openai", ) - # Initialize OpenAI client - client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY", "sk-test")) + # Initialize OpenAI client wired to mock-llm + client = mock_openai_client() @trace(event_type="model") def call_openai(prompt: str) -> str: """Call OpenAI and enrich span.""" try: response = client.chat.completions.create( - model="gpt-3.5-turbo", + model=MOCK_LLM_MODEL, messages=[{"role": "user", "content": prompt}], max_tokens=10, ) result = response.choices[0].message.content or "" - # Enrich with OpenAI metadata + # Enrich with model metadata tracer.enrich_span( metadata={ - "model": "gpt-3.5-turbo", + "model": MOCK_LLM_MODEL, "prompt": prompt, "response": result, }, @@ -199,13 +195,11 @@ def call_openai(prompt: str) -> str: return result except Exception as e: - tracer.enrich_span(metadata={"error": str(e), "model": "gpt-3.5-turbo"}) + tracer.enrich_span(metadata={"error": str(e), "model": MOCK_LLM_MODEL}) raise - # Execute (will skip if no OPENAI_API_KEY) - if os.environ.get("OPENAI_API_KEY"): - result = call_openai("Say 'test'") - assert isinstance(result, str) + result = call_openai("Say 'test'") + assert isinstance(result, str) @pytest.mark.integration diff --git a/tests/integration/test_evaluate_enrich.py b/tests/integration/test_evaluate_enrich.py deleted file mode 100644 index d6d3300a..00000000 --- a/tests/integration/test_evaluate_enrich.py +++ /dev/null @@ -1,254 +0,0 @@ -"""Integration tests for evaluate() + enrich_span() pattern. - -⚠️ SKIPPED: Pending v1 evaluation API migration -This test suite is skipped because the evaluate() function no longer exists in v1. -The v1 evaluation API uses a different pattern and these tests need to be migrated. - -This module tests the end-to-end functionality of the evaluate() pattern -with enrich_span() calls, validating that tracer discovery works correctly -via baggage propagation after the v1.0 selective propagation fix. -""" - -# pylint: disable=unused-argument,import-outside-toplevel,unused-variable,unused-import -# Justification: Dynamic test imports; unused: Test scaffolding -# Justification: -# - Unused datapoint arg in test fixture - -import os -from typing import Any, Dict - -import pytest - -# Skip entire module - v0 evaluate() function no longer exists in v1 -pytestmark = pytest.mark.skip( - reason="Skipped pending v1 evaluation API migration - evaluate() function no longer exists in v1" -) - -# Import handling: evaluate() doesn't exist in v1, but we keep the import -# for reference. The module is skipped so tests won't run anyway. -try: - from honeyhive import HoneyHiveTracer, enrich_span, evaluate -except ImportError: - # evaluate() doesn't exist in v1 - this is expected - # Module is skipped via pytestmark above - HoneyHiveTracer = None # type: ignore - enrich_span = None # type: ignore - evaluate = None # type: ignore - - -@pytest.mark.integration -@pytest.mark.skipif( - not os.environ.get("HH_API_KEY"), reason="Requires HH_API_KEY environment variable" -) -class TestEvaluateEnrichIntegration: - """Test evaluate() with enrich_span() pattern (v1.0 baggage fix validation).""" - - def test_evaluate_with_enrich_span_tracer_discovery(self) -> None: - """Test that enrich_span() works within evaluate() via tracer discovery. - - This test validates the v1.0 fix for selective baggage propagation. - The tracer should be discovered via honeyhive_tracer_id in baggage. - """ - # Track calls - calls: list = [] - - def user_function(datapoint: Dict[str, Any]) -> Dict[str, Any]: - """User function with enrich_span call.""" - calls.append("function_called") - - # This should work via tracer discovery (v1.0 fix) - enrich_span( - metadata={"input": datapoint["inputs"]}, - metrics={"call_count": len(calls)}, - ) - calls.append("enrich_called") - - return {"output": "test_result", "status": "success"} - - # Run evaluation with small dataset - result = evaluate( - function=user_function, - dataset=[{"inputs": {"text": "test1"}}, {"inputs": {"text": "test2"}}], - api_key=os.environ["HH_API_KEY"], - project="test-evaluate-enrich-integration", - name="v1.0-baggage-fix-validation", - ) - - # Verify evaluation completed - assert result is not None - assert hasattr(result, "status") - - # Verify both datapoints were processed - assert len(calls) >= 4 # 2 function calls + 2 enrich calls - - def test_evaluate_with_explicit_tracer_enrich(self) -> None: - """Test evaluate() with explicit tracer and instance method enrichment. - - This is the recommended pattern for v1.0+ (instance methods). - """ - tracer = HoneyHiveTracer( - api_key=os.environ["HH_API_KEY"], - project="test-evaluate-explicit-tracer", - session_name="v1.0-instance-method-pattern", - ) - - def user_function_with_tracer(datapoint: Dict[str, Any]) -> Dict[str, Any]: - """User function with explicit tracer (recommended pattern).""" - # Instance method pattern (PRIMARY in v1.0) - tracer.enrich_span( - metadata={"input": datapoint["inputs"]}, - metrics={"datapoint_processed": 1}, - ) - - return {"output": "processed", "status": "success"} - - # Run evaluation - result = evaluate( - function=user_function_with_tracer, - dataset=[{"inputs": {"text": "test"}}], - api_key=os.environ["HH_API_KEY"], - project="test-evaluate-explicit-tracer", - name="explicit-tracer-pattern", - ) - - assert result is not None - assert hasattr(result, "status") - - def test_evaluate_enrich_span_with_evaluation_context(self) -> None: - """Test that evaluation context (run_id, datapoint_id) propagates correctly. - - Validates that the v1.0 selective baggage fix propagates - evaluation context keys (run_id, dataset_id, datapoint_id). - """ - captured_metadata: list = [] - - def user_function(datapoint: Dict[str, Any]) -> Dict[str, Any]: - """Capture metadata to verify context propagation.""" - # Enrich with metadata - enrich_span( - metadata={ - "datapoint_input": datapoint["inputs"], - "test_marker": "context_propagation_test", - } - ) - - captured_metadata.append(datapoint["inputs"]) - return {"output": "ok"} - - # Run evaluation - result = evaluate( - function=user_function, - dataset=[ - {"inputs": {"idx": 1}}, - {"inputs": {"idx": 2}}, - {"inputs": {"idx": 3}}, - ], - api_key=os.environ["HH_API_KEY"], - project="test-evaluation-context-propagation", - name="context-propagation-validation", - ) - - # Verify all datapoints processed - assert len(captured_metadata) == 3 - assert result is not None - - def test_evaluate_child_spans_have_evaluation_metadata(self) -> None: - """Test that child spans created during evaluate() have evaluation metadata. - - This test validates the baggage propagation fix that ensures run_id, - dataset_id, and datapoint_id propagate to all child spans. - """ - import time - - from honeyhive import HoneyHive, trace - - # Track span creation - span_names = [] - - @trace(event_type="tool", event_name="child_operation") - def child_operation(text: str) -> str: - """Child function that creates a span.""" - span_names.append("child_operation") - return text.upper() - - def user_function(datapoint: Dict[str, Any]) -> Dict[str, Any]: - """Function that creates child spans.""" - inputs = datapoint.get("inputs", {}) - text = inputs.get("text", "") - - # Create child span - result = child_operation(text) - - return {"output": result, "status": "success"} - - # Run evaluation - result = evaluate( - function=user_function, - dataset=[ - {"inputs": {"text": "test1"}}, - {"inputs": {"text": "test2"}}, - ], - api_key=os.environ["HH_API_KEY"], - project="test-evaluation-metadata-propagation", - name="child-span-metadata-test", - ) - - # Verify evaluation completed - assert result is not None - assert hasattr(result, "status") - assert result.status == "completed" - - # Verify child spans were created - assert len(span_names) == 2 # One child span per datapoint - - # Give backend time to process spans - time.sleep(3) - - # Verify evaluation metadata was set - assert hasattr(result, "run_id") - run_id = result.run_id - - # Validate that run was created with correct structure - # The backend validation happens during evaluate() execution - # If child spans have evaluation metadata, the run linking will work correctly - - # NOTE: Full backend validation would require: - # 1. Fetching the run via API - # 2. Fetching associated events/sessions - # 3. Validating run_id, dataset_id, datapoint_id in event metadata - # - # This is tested implicitly by the evaluate() success and the verbose - # logs showing the attributes are set on spans before export. - - def test_evaluate_enrich_span_error_handling(self) -> None: - """Test that enrich_span gracefully handles errors in evaluate(). - - Validates that enrichment failures don't crash evaluation. - """ - processed_count = 0 - - def user_function_with_error(datapoint: Dict[str, Any]) -> Dict[str, Any]: - """Function that attempts enrichment.""" - nonlocal processed_count - processed_count += 1 - - # This might fail but shouldn't crash - try: - enrich_span(metadata={"count": processed_count}) - except Exception: - pass # Graceful degradation - - return {"output": f"processed_{processed_count}"} - - # Run evaluation - result = evaluate( - function=user_function_with_error, - dataset=[{"inputs": {"test": i}} for i in range(5)], - api_key=os.environ["HH_API_KEY"], - project="test-enrich-error-handling", - name="error-handling-validation", - ) - - # Should complete despite any enrichment issues - assert processed_count == 5 - assert result is not None diff --git a/tests/integration/test_openai_integration.py b/tests/integration/test_openai_integration.py index 14db13eb..83a55ec6 100644 --- a/tests/integration/test_openai_integration.py +++ b/tests/integration/test_openai_integration.py @@ -12,7 +12,6 @@ Environment Variables: HH_API_KEY: HoneyHive API key HH_PROJECT: HoneyHive project name - OPENAI_API_KEY: OpenAI API key LKGV (last known good versions) for this path are pinned in pyproject.toml. """ @@ -22,13 +21,14 @@ import pytest -# Skip entire module if keys not present +from tests.integration._mock_llm_helpers import ( + MOCK_LLM_MODEL, + MOCK_RESPONSE, + mock_openai_client, +) + pytestmark = [ pytest.mark.skipif(not os.getenv("HH_API_KEY"), reason="HH_API_KEY not set"), - pytest.mark.skipif( - not os.getenv("OPENAI_API_KEY"), reason="OPENAI_API_KEY not set" - ), - pytest.mark.openai, pytest.mark.slow, ] @@ -44,7 +44,6 @@ def setup(self): def test_basic_chat_completion(self): """Test basic chat completion is traced correctly.""" - import openai from openinference.instrumentation.openai import OpenAIInstrumentor from honeyhive import HoneyHiveTracer @@ -62,9 +61,9 @@ def test_basic_chat_completion(self): try: # Make OpenAI call - client = openai.OpenAI() + client = mock_openai_client() response = client.chat.completions.create( - model="gpt-4o-mini", + model=MOCK_LLM_MODEL, messages=[{"role": "user", "content": "Say 'test' and nothing else."}], max_tokens=10, ) @@ -83,7 +82,6 @@ def test_basic_chat_completion(self): def test_chat_completion_with_enrichment(self): """Test that enrich_span works within OpenAI traced calls.""" - import openai from openinference.instrumentation.openai import OpenAIInstrumentor from honeyhive import HoneyHiveTracer, enrich_span, trace @@ -104,9 +102,9 @@ def process_with_openai(prompt: str) -> str: """Process a prompt with OpenAI and enrich the span.""" enrich_span(metadata={"prompt_length": len(prompt)}) - client = openai.OpenAI() + client = mock_openai_client() response = client.chat.completions.create( - model="gpt-4o-mini", + model=MOCK_LLM_MODEL, messages=[{"role": "user", "content": prompt}], max_tokens=20, ) @@ -129,7 +127,6 @@ def process_with_openai(prompt: str) -> str: def test_streaming_completion(self): """Test streaming chat completion is traced.""" - import openai from openinference.instrumentation.openai import OpenAIInstrumentor from honeyhive import HoneyHiveTracer @@ -144,9 +141,9 @@ def test_streaming_completion(self): instrumentor.instrument(tracer_provider=tracer.provider) try: - client = openai.OpenAI() + client = mock_openai_client() stream = client.chat.completions.create( - model="gpt-4o-mini", + model=MOCK_LLM_MODEL, messages=[{"role": "user", "content": "Count from 1 to 5."}], max_tokens=50, stream=True, @@ -157,9 +154,13 @@ def test_streaming_completion(self): if chunk.choices[0].delta.content: chunks.append(chunk.choices[0].delta.content) + # Verify the stream actually chunked rather than buffering the full + # response into one delta, AND that concatenating deltas exactly + # reconstructs MOCK_RESPONSE. The exact-match assertion catches + # both "stream collapsed into one delta" and content corruption. full_response = "".join(chunks) - assert len(full_response) > 0 - assert any(str(i) in full_response for i in range(1, 6)) + assert len(chunks) > 1 + assert full_response == MOCK_RESPONSE tracer.flush() @@ -178,7 +179,6 @@ def setup(self): def test_basic_chat_completion_traceloop(self): """Test basic chat completion with Traceloop instrumentor.""" - import openai from opentelemetry.instrumentation.openai import OpenAIInstrumentor from honeyhive import HoneyHiveTracer @@ -193,9 +193,9 @@ def test_basic_chat_completion_traceloop(self): instrumentor.instrument(tracer_provider=tracer.provider) try: - client = openai.OpenAI() + client = mock_openai_client() response = client.chat.completions.create( - model="gpt-4o-mini", + model=MOCK_LLM_MODEL, messages=[ { "role": "user", @@ -215,7 +215,6 @@ def test_basic_chat_completion_traceloop(self): def test_nested_traces_with_openai(self): """Test nested @trace decorators with OpenAI calls.""" - import openai from opentelemetry.instrumentation.openai import OpenAIInstrumentor from honeyhive import HoneyHiveTracer, trace @@ -240,9 +239,9 @@ def outer_function(query: str) -> Dict[str, Any]: @trace(event_type="tool") def inner_function(text: str) -> str: """Inner traced function that calls OpenAI.""" - client = openai.OpenAI() + client = mock_openai_client() response = client.chat.completions.create( - model="gpt-4o-mini", + model=MOCK_LLM_MODEL, messages=[{"role": "user", "content": f"Summarize: {text}"}], max_tokens=30, ) diff --git a/tests/integration/test_otel_context_propagation_integration.py b/tests/integration/test_otel_context_propagation_integration.py index 94f0db4a..0509729d 100644 --- a/tests/integration/test_otel_context_propagation_integration.py +++ b/tests/integration/test_otel_context_propagation_integration.py @@ -72,7 +72,8 @@ def test_w3c_trace_context_injection_extraction( assert parts[0] == "00" # version assert len(parts[1]) == 32 # trace_id (128-bit hex) assert len(parts[2]) == 16 # span_id (64-bit hex) - assert parts[3] in ["00", "01"] # trace_flags + assert len(parts[3]) == 2 # trace_flags (2-char hex per W3C spec) + int(parts[3], 16) # must be valid hex # Extract trace context from carrier (simulates incoming HTTP request) extracted_context = propagator.extract(carrier) diff --git a/tests/integration/test_tracing_integration.py b/tests/integration/test_tracing_integration.py index 50d6b0b6..c1ab1244 100644 --- a/tests/integration/test_tracing_integration.py +++ b/tests/integration/test_tracing_integration.py @@ -18,6 +18,8 @@ import pytest +from tests.integration._mock_llm_helpers import MOCK_LLM_MODEL, mock_openai_client + # Skip entire module if key not present pytestmark = [ pytest.mark.skipif(not os.getenv("HH_API_KEY"), reason="HH_API_KEY not set"), @@ -817,13 +819,7 @@ def test_openai_inputs_outputs_verification(self, fetch_events): """ import time - # Skip if OpenAI not available - openai_key = os.getenv("OPENAI_API_KEY") - if not openai_key: - pytest.skip("OPENAI_API_KEY not set") - try: - import openai from openinference.instrumentation.openai import OpenAIInstrumentor except ImportError: pytest.skip("openai or openinference not installed") @@ -842,11 +838,11 @@ def test_openai_inputs_outputs_verification(self, fetch_events): instrumentor.instrument(tracer_provider=tracer.provider) try: - client = openai.OpenAI() + client = mock_openai_client() # Use a unique prompt we can verify was captured test_prompt = "Say exactly: 'integration test verification'" - test_model = "gpt-3.5-turbo" + test_model = MOCK_LLM_MODEL response = client.chat.completions.create( model=test_model, diff --git a/tests/unit/test_tracer_processing_otlp_exporter.py b/tests/unit/test_tracer_processing_otlp_exporter.py index 718fe496..63baba5c 100644 --- a/tests/unit/test_tracer_processing_otlp_exporter.py +++ b/tests/unit/test_tracer_processing_otlp_exporter.py @@ -1008,16 +1008,37 @@ class TestOTLPJSONExporterAnyValueMapping: def test_string_attribute_preserved_as_string(self) -> None: assert OTLPJSONExporter._to_otlp_any_value("hello") == {"stringValue": "hello"} - def test_int_attribute_preserved_as_int(self) -> None: - # Integers must serialize as intValue (not stringified) so the backend - # stores them as numbers — this is the primary bug from HHAI-4935. - assert OTLPJSONExporter._to_otlp_any_value(42) == {"intValue": 42} - - def test_large_int_preserved_as_int(self) -> None: + def test_int_attribute_serialized_as_string(self) -> None: + # Integers must serialize as intValue JSON strings per protobuf JSON + # mapping spec. Raw JSON numbers lose precision above 2^53 through the + # server's float64 decode path (HHAI-5004). + assert OTLPJSONExporter._to_otlp_any_value(42) == {"intValue": "42"} + + def test_large_int_preserved_exactly_as_string(self) -> None: + # int64 max — must round-trip exactly; a raw JSON number would be fine + # here but values just above 2^53 would not. assert OTLPJSONExporter._to_otlp_any_value(9223372036854775807) == { - "intValue": 9223372036854775807 + "intValue": "9223372036854775807" } + def test_int_above_float64_precision_round_trips_exactly(self) -> None: + # 2^53 + 1 = 9007199254740993 cannot be represented exactly as float64. + # Emitting it as a JSON string ensures the server recovers the exact value. + import json as _json + + large = 2**53 + 1 # 9007199254740993 + result = OTLPJSONExporter._to_otlp_any_value(large) + assert result == {"intValue": "9007199254740993"} + # Prove the wire JSON encodes the value as a quoted string, not a bare + # number. json.loads returns str for JSON strings and int for JSON numbers, + # so this assertion fails if str() is ever dropped from the helper. + wire = _json.dumps(result) + recovered = _json.loads(wire) + assert isinstance(recovered["intValue"], str), ( + "intValue must be a JSON string on the wire, not a number" + ) + assert int(recovered["intValue"]) == large + def test_float_attribute_preserved_as_double(self) -> None: assert OTLPJSONExporter._to_otlp_any_value(3.14) == {"doubleValue": 3.14} @@ -1058,9 +1079,9 @@ def test_list_attribute_recurses(self) -> None: assert OTLPJSONExporter._to_otlp_any_value([1, 2, 3]) == { "arrayValue": { "values": [ - {"intValue": 1}, - {"intValue": 2}, - {"intValue": 3}, + {"intValue": "1"}, + {"intValue": "2"}, + {"intValue": "3"}, ] } } @@ -1094,7 +1115,7 @@ def test_key_values_builder_preserves_types(self) -> None: ) by_key = {kv["key"]: kv["value"] for kv in result} assert by_key == { - "count": {"intValue": 42}, + "count": {"intValue": "42"}, "name": {"stringValue": "foo"}, "ok": {"boolValue": True}, "ratio": {"doubleValue": 0.5}, @@ -1155,20 +1176,73 @@ def test_span_attributes_emit_typed_any_values( } assert resource_attrs == { "service.name": {"stringValue": "svc"}, - "replicas": {"intValue": 3}, + "replicas": {"intValue": "3"}, } # Span attributes preserve native types (the HHAI-4935 regression). span_payload = resource_span["scopeSpans"][0]["spans"][0] span_attrs = {kv["key"]: kv["value"] for kv in span_payload["attributes"]} assert span_attrs == { - "attr.int": {"intValue": 42}, + "attr.int": {"intValue": "42"}, "attr.float": {"doubleValue": 3.14}, "attr.bool": {"boolValue": True}, "attr.string": {"stringValue": "hello"}, } +class TestOTLPJSONExporterTimestamps: + """Verify uint64 timestamp fields are serialized as JSON strings (HHAI-5004).""" + + @patch("honeyhive.tracer.processing.otlp_exporter.requests.Session") + def test_span_timestamps_are_strings( + self, mock_session_class: Mock, mock_tracer: Mock + ) -> None: + mock_session = Mock() + mock_response = Mock() + mock_response.status_code = 200 + mock_response.text = "" + mock_session.post.return_value = mock_response + mock_session_class.return_value = mock_session + + span = Mock(spec=ReadableSpan) + span.name = "ts_test" + span.context = Mock() + span.context.trace_id = 0x1234567890ABCDEF1234567890ABCDEF + span.context.span_id = 0x1234567890ABCDEF + span.parent = None + span.kind = Mock() + span.kind.name = "INTERNAL" + span.start_time = 1_000_000_000 + span.end_time = 2_000_000_000 + span.status = Mock() + span.status.status_code = Mock() + span.status.status_code.name = "UNSET" + span.status.description = None + span.attributes = {} + span.resource = Mock() + span.resource.attributes = {} + span.instrumentation_scope = None + + event = Mock() + event.timestamp = 1_500_000_000 + event.name = "evt" + event.attributes = {} + span.events = [event] + + exporter = OTLPJSONExporter(TEST_OTLP_ENDPOINT, tracer_instance=mock_tracer) + exporter.export([span]) + + import json as _json + + payload = _json.loads(mock_session.post.call_args[1]["data"]) + span_json = payload["resourceSpans"][0]["scopeSpans"][0]["spans"][0] + + # protobuf JSON mapping requires uint64 fields to be JSON strings + assert span_json["startTimeUnixNano"] == "1000000000" + assert span_json["endTimeUnixNano"] == "2000000000" + assert span_json["events"][0]["timeUnixNano"] == "1500000000" + + class TestHoneyHiveOTLPExporterProtocol: """Test HoneyHive OTLP exporter protocol selection."""