@@ -71,6 +71,7 @@ def extract(
7171 * ,
7272 schema : Dict [str , object ],
7373 url : str ,
74+ actions : Iterable [web_extract_params .Action ] | Omit = omit ,
7475 fact_check : bool | Omit = omit ,
7576 follow_subdomains : bool | Omit = omit ,
7677 include_frames : bool | Omit = omit ,
@@ -96,13 +97,20 @@ def extract(
9697 relevant internal links, and extract structured data from the selected pages.
9798
9899 Args:
99- schema: JSON Schema for the returned data object. TypeScript Zod users can pass a JSON
100- Schema generated from a Zod object; Python users can pass the equivalent JSON
101- Schema object.
100+ schema: JSON Schema for the returned data object. Image fields such as `image_urls` or
101+ `product_photos` automatically make page image references available to
102+ extraction, so product data and photos can be returned in one call. TypeScript
103+ Zod users can pass a JSON Schema generated from a Zod object; Python users can
104+ pass the equivalent JSON Schema object.
102105
103106 url: The starting website URL to crawl and extract from. Must include http:// or
104107 https://.
105108
109+ actions: Optional browser actions executed in order on the requested page after it loads,
110+ before links are discovered or additional pages are crawled. Requires a paid
111+ plan. When actions are provided and stopAfterMs is omitted, the crawl budget
112+ defaults to 110000 ms.
113+
106114 fact_check: When true, every returned value must be grounded in facts stated on the page;
107115 fields that cannot be supported by the page are returned as null/empty. When
108116 false (default), the model may make reasonable inferences and derivations from
@@ -130,7 +138,8 @@ def extract(
130138 exchange for more stable output on animated pages.
131139
132140 stop_after_ms: Soft time budget for the crawl in milliseconds. Min: 10000 (10s). Max: 110000
133- (110s). Default: 80000 (80s).
141+ (110s). Defaults to 80000 (80s), or 110000 (110s) when browser actions are
142+ provided.
134143
135144 tags: Optional tags for tracking usage. Up to 20 tags, each 1 to 50 characters.
136145
@@ -155,6 +164,7 @@ def extract(
155164 {
156165 "schema" : schema ,
157166 "url" : url ,
167+ "actions" : actions ,
158168 "fact_check" : fact_check ,
159169 "follow_subdomains" : follow_subdomains ,
160170 "include_frames" : include_frames ,
@@ -1338,13 +1348,15 @@ def web_crawl_md(
13381348 than this value, it will be aborted with a 408 status code. Maximum allowed
13391349 value is 300000ms (5 minutes).
13401350
1341- url_regex: Regex pattern. Only URLs matching this pattern will be followed and scraped.
1351+ url_regex: Regex pattern. Only URLs matching this pattern will be followed and scraped. An
1352+ automatic prefix scope in the form ^<starting URL> follows a redirect of the
1353+ starting page.
13421354
13431355 use_main_content_only: Extract only the main content, stripping headers, footers, sidebars, and
13441356 navigation
13451357
1346- wait_for_ms: Optional browser wait time in milliseconds after initial page load for each
1347- crawled page . Min: 0. Max: 30000 (30 seconds).
1358+ wait_for_ms: Browser wait time in milliseconds after initial page load for each crawled page.
1359+ Defaults to 3500 (3.5 seconds) . Min: 0. Max: 30000 (30 seconds).
13481360
13491361 zdr: Set to enabled to bypass shared caches and omit request and response content
13501362 from retained usage logs. Requires zero data retention to be enabled for your
@@ -2309,6 +2321,7 @@ async def extract(
23092321 * ,
23102322 schema : Dict [str , object ],
23112323 url : str ,
2324+ actions : Iterable [web_extract_params .Action ] | Omit = omit ,
23122325 fact_check : bool | Omit = omit ,
23132326 follow_subdomains : bool | Omit = omit ,
23142327 include_frames : bool | Omit = omit ,
@@ -2334,13 +2347,20 @@ async def extract(
23342347 relevant internal links, and extract structured data from the selected pages.
23352348
23362349 Args:
2337- schema: JSON Schema for the returned data object. TypeScript Zod users can pass a JSON
2338- Schema generated from a Zod object; Python users can pass the equivalent JSON
2339- Schema object.
2350+ schema: JSON Schema for the returned data object. Image fields such as `image_urls` or
2351+ `product_photos` automatically make page image references available to
2352+ extraction, so product data and photos can be returned in one call. TypeScript
2353+ Zod users can pass a JSON Schema generated from a Zod object; Python users can
2354+ pass the equivalent JSON Schema object.
23402355
23412356 url: The starting website URL to crawl and extract from. Must include http:// or
23422357 https://.
23432358
2359+ actions: Optional browser actions executed in order on the requested page after it loads,
2360+ before links are discovered or additional pages are crawled. Requires a paid
2361+ plan. When actions are provided and stopAfterMs is omitted, the crawl budget
2362+ defaults to 110000 ms.
2363+
23442364 fact_check: When true, every returned value must be grounded in facts stated on the page;
23452365 fields that cannot be supported by the page are returned as null/empty. When
23462366 false (default), the model may make reasonable inferences and derivations from
@@ -2368,7 +2388,8 @@ async def extract(
23682388 exchange for more stable output on animated pages.
23692389
23702390 stop_after_ms: Soft time budget for the crawl in milliseconds. Min: 10000 (10s). Max: 110000
2371- (110s). Default: 80000 (80s).
2391+ (110s). Defaults to 80000 (80s), or 110000 (110s) when browser actions are
2392+ provided.
23722393
23732394 tags: Optional tags for tracking usage. Up to 20 tags, each 1 to 50 characters.
23742395
@@ -2393,6 +2414,7 @@ async def extract(
23932414 {
23942415 "schema" : schema ,
23952416 "url" : url ,
2417+ "actions" : actions ,
23962418 "fact_check" : fact_check ,
23972419 "follow_subdomains" : follow_subdomains ,
23982420 "include_frames" : include_frames ,
@@ -3576,13 +3598,15 @@ async def web_crawl_md(
35763598 than this value, it will be aborted with a 408 status code. Maximum allowed
35773599 value is 300000ms (5 minutes).
35783600
3579- url_regex: Regex pattern. Only URLs matching this pattern will be followed and scraped.
3601+ url_regex: Regex pattern. Only URLs matching this pattern will be followed and scraped. An
3602+ automatic prefix scope in the form ^<starting URL> follows a redirect of the
3603+ starting page.
35803604
35813605 use_main_content_only: Extract only the main content, stripping headers, footers, sidebars, and
35823606 navigation
35833607
3584- wait_for_ms: Optional browser wait time in milliseconds after initial page load for each
3585- crawled page . Min: 0. Max: 30000 (30 seconds).
3608+ wait_for_ms: Browser wait time in milliseconds after initial page load for each crawled page.
3609+ Defaults to 3500 (3.5 seconds) . Min: 0. Max: 30000 (30 seconds).
35863610
35873611 zdr: Set to enabled to bypass shared caches and omit request and response content
35883612 from retained usage logs. Requires zero data retention to be enabled for your
0 commit comments