Fix turn_taking metric failing when agent_perf_stats.csv prompt field is large - #218
Open
fanny-riols wants to merge 1 commit into
Open
Fix turn_taking metric failing when agent_perf_stats.csv prompt field is large#218fanny-riols wants to merge 1 commit into
fanny-riols wants to merge 1 commit into
Conversation
… is large mean_agent_perf_stat() used csv.DictReader with Python's default 131072-byte field size limit. Agents with large system prompts (e.g. medical_hr's ~40-tool config) push the serialized prompt field in agent_perf_stats.csv past that limit, raising _csv.Error. Since turn_taking.compute() calls mean_agent_perf_stat() unguarded, that error propagated up and zeroed out the entire turn_taking metric (score forced to 0, details wiped) instead of just the token sub-metrics that depend on this file. - Raise csv.field_size_limit() well above any observed field size. - Also catch csv.Error in mean_agent_perf_stat() and degrade to None, so a pathologically large field can never again take down the whole metric. - Add regression tests.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Issue: prompt too big and reading from the csv failed, so the turn taking metric wasn't computed.
More details:
mean_agent_perf_stat()used csv.DictReader with Python's default 131072-byte field size limit. Agents with large system prompts (e.g. medical_hr's ~40-tool config) push the serialized prompt field in agent_perf_stats.csv past that limit, raising _csv.Error. Sinceturn_taking.compute()calls mean_agent_perf_stat() unguarded, that error propagated up and zeroed out the entire turn_taking metric (score forced to 0, details wiped) instead of just the token sub-metrics that depend on this file.I tested that change and it fixed my problem.