Conversation
A Swift Testing test that records several failed #expect assertions was counted as several failed tests. The run summary reports issues rather than failed tests, and the run state counted failure diagnostics as failed tests, so progress and the final summary could report more failures than tests run and zero passes. Count native failed result lines per test, attribute their issue counts, and only let unattributed summary issues mark unreported tests as failed. Count failure diagnostics by distinct test when reconciling the summary. Fixes getsentry#533 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #533
Problem
When one Swift Testing test records several failed
#expectassertions, progress and the test summary count each assertion as a failed test. With the input from #533 (one passing test, one failing test with 3 issues), the stream reports:completed: 2, failed: 3totalTests: 3, passedTests: 0, failedTests: 3The reporter's real run of 16 tests (8 passed, 8 failed) came out as
totalTests: 25, passedTests: 0, failedTests: 25. A consumer that checks progress against the test-case results rejects that stream as contradictory.There are two causes:
createXcodebuildEventParserskips native✘ Test "..." failedlines when counting failures and uses the run summary'swith N issuescount instead.createTestSummaryFragment(andcreateStateTestCountsin the domain results) takeMath.max(failedTests, testFailures.length), andtestFailuresholds one entry per diagnostic, not per test.Fix
with N issues). At the run summary, it subtracts the issues already attributed to reported failures. Any issues left over can mark unreported tests as failed, capped at the number of unreported tests. This keeps the behavior fromdefers Swift Testing failure progress until the run summaryfor summary-only failures.countFailedTestshelper inxcodebuild-run-state.tscounts failure diagnostics by distinct test. The run-state summary and the domain-result counts both use it.For the #533 input, the stream now reports progress
completed: 2, failed: 1and summarytotalTests: 2, passedTests: 1, failedTests: 1. All three failure diagnostics are still emitted.Tests
xcodebuild-event-parser.test.ts:counts a Swift Testing test with several issues as one failed testreplays the [Bug]: Swift Testing JSONL counts assertion issues as failed tests #533 input through the parser and run state.xcodebuild-run-state.test.ts:counts several failure diagnostics from one test as one failed test.Both tests fail on
d13ff0c(2 failed / 50 passed in the two files) and pass with the fix (52/52). The full unit suite passes (243 files, 2593 tests).npm run build,npm run typecheck,npm run lint(0 errors) andprettier --checkon the changed files also pass. I did not run the snapshot and smoke suites.I wrote this change with AI assistance (Claude) and checked it against the reproduction in #533 and the tests above.