@@ -86,10 +86,8 @@ def extract(
8686 timeout : float | httpx .Timeout | None | NotGiven = not_given ,
8787 ) -> WebExtractResponse :
8888 """
89- Crawl a website, convert pages to Markdown using the scrape cache, and extract
90- structured data into the provided JSON Schema. The schema must describe the
91- response data object. This endpoint does not accept targeted page-type
92- selection.
89+ Crawl a website, use the provided JSON Schema and instructions to prioritize
90+ relevant internal links, and extract structured data from the selected pages.
9391
9492 Args:
9593 schema: JSON Schema for the returned data object. TypeScript Zod users can pass a JSON
@@ -99,12 +97,12 @@ def extract(
9997 url: The starting website URL to crawl and extract from. Must include http:// or
10098 https://.
10199
102- fact_check: When true (default) , every returned value must be grounded in facts stated on
103- the page; fields that cannot be supported by the page are returned as
104- null/empty. When false, the model may make reasonable inferences and derivations
105- from the page content (e.g. ideal customer, competitor analysis,
106- recommendations) while keeping verifiable specifics (names, quotes, URLs, dates,
107- metrics) faithful to the source.
100+ fact_check: When true, every returned value must be grounded in facts stated on the page;
101+ fields that cannot be supported by the page are returned as null/empty. When
102+ false (default) , the model may make reasonable inferences and derivations from
103+ the page content (e.g. ideal customer, competitor analysis, recommendations)
104+ while keeping verifiable specifics (names, quotes, URLs, dates, metrics)
105+ faithful to the source.
108106
109107 follow_subdomains: When true, follow links on subdomains of the starting URL's domain.
110108
@@ -114,7 +112,7 @@ def extract(
114112 interpret fields in the schema.
115113
116114 max_age_ms: Return cached scrape results if a prior scrape for the same parameters is
117- younger than this many milliseconds.
115+ younger than this many milliseconds. Defaults to 7 days (604800000 ms).
118116
119117 stop_after_ms: Soft time budget for the crawl in milliseconds.
120118
@@ -876,10 +874,8 @@ async def extract(
876874 timeout : float | httpx .Timeout | None | NotGiven = not_given ,
877875 ) -> WebExtractResponse :
878876 """
879- Crawl a website, convert pages to Markdown using the scrape cache, and extract
880- structured data into the provided JSON Schema. The schema must describe the
881- response data object. This endpoint does not accept targeted page-type
882- selection.
877+ Crawl a website, use the provided JSON Schema and instructions to prioritize
878+ relevant internal links, and extract structured data from the selected pages.
883879
884880 Args:
885881 schema: JSON Schema for the returned data object. TypeScript Zod users can pass a JSON
@@ -889,12 +885,12 @@ async def extract(
889885 url: The starting website URL to crawl and extract from. Must include http:// or
890886 https://.
891887
892- fact_check: When true (default) , every returned value must be grounded in facts stated on
893- the page; fields that cannot be supported by the page are returned as
894- null/empty. When false, the model may make reasonable inferences and derivations
895- from the page content (e.g. ideal customer, competitor analysis,
896- recommendations) while keeping verifiable specifics (names, quotes, URLs, dates,
897- metrics) faithful to the source.
888+ fact_check: When true, every returned value must be grounded in facts stated on the page;
889+ fields that cannot be supported by the page are returned as null/empty. When
890+ false (default) , the model may make reasonable inferences and derivations from
891+ the page content (e.g. ideal customer, competitor analysis, recommendations)
892+ while keeping verifiable specifics (names, quotes, URLs, dates, metrics)
893+ faithful to the source.
898894
899895 follow_subdomains: When true, follow links on subdomains of the starting URL's domain.
900896
@@ -904,7 +900,7 @@ async def extract(
904900 interpret fields in the schema.
905901
906902 max_age_ms: Return cached scrape results if a prior scrape for the same parameters is
907- younger than this many milliseconds.
903+ younger than this many milliseconds. Defaults to 7 days (604800000 ms).
908904
909905 stop_after_ms: Soft time budget for the crawl in milliseconds.
910906
0 commit comments