Skip to content

refactor: unify VLM on deepseek-flash and tighten page-text handling - #442

Merged
EricNGOntos merged 2 commits into
mainfrom
refactor/wuchengke/vlm-and-page-text
Sep 28, 2026
Merged

EricNGOntos merged 2 commits into
mainfrom
refactor/wuchengke/vlm-and-page-text

Conversation

@EricNGOntos

Copy link
Copy Markdown
Contributor

Summary

  • Switch the default vision model to official deepseek-flash, drop IMAGE_MODEL_MAX, and let ASSET_MODEL follow IMAGE_MODEL unless overridden.
  • Tighten PROFILE page-text bands and TOC-anchor selection, and keep Stage-1 debug TOC dumps split by region.

Test plan

  • make check
  • Focused worker/shared/ECS tests for page text, TOC anchors, asset model, and task-definition env

Made with Cursor

EricNGOntos and others added 2 commits September 28, 2026 17:28
…page text handling

- Consolidated VLM model references to use `deepseek-flash` across documentation and configuration files.
- Enhanced the `PageTextBands` structure to support ordered line records, improving text extraction and processing.
- Removed deprecated fields and streamlined the handling of page text in various components, ensuring consistency in data representation.
- Updated tests to reflect changes in the `PageTextBands` structure and VLM model usage.
The previous dump flattened every extracted TOC into one tree, which hid later regions in multi-TOC documents.

Co-authored-by: Cursor <cursoragent@cursor.com>
@EricNGOntos
EricNGOntos merged commit a249e0f into main Sep 28, 2026
6 checks passed
@EricNGOntos
EricNGOntos deleted the refactor/wuchengke/vlm-and-page-text branch September 28, 2026 09:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant