> ## Documentation Index
> Fetch the complete documentation index at: https://docs.geekflare.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Web Scraping

> Fetch a page and return content as Markdown, HTML, JSON, or plain text. Automatically detects whether JavaScript rendering is needed, with optional stealth mode, proxy routing (`proxyCountry`, with optional `proxyMode` control), CSS/XPath field extraction, and ready-made `product`/`contact` extraction templates.



## OpenAPI

````yaml POST /webscraping
openapi: 3.1.0
info:
  title: Geekflare
  description: Official OpenAPI specification for all Geekflare endpoints.
  version: 1.0.0
  license:
    name: MIT
servers:
  - url: https://api.geekflare.com
security:
  - x-api-key: []
paths:
  /webscraping:
    post:
      tags:
        - api-tool
      summary: Scrape a webpage with custom options
      description: >-
        Fetch a page and return content as Markdown, HTML, JSON, or plain text.
        Automatically detects whether JavaScript rendering is needed, with
        optional stealth mode, proxy routing (`proxyCountry`, with optional
        `proxyMode` control), CSS/XPath field extraction, and ready-made
        `product`/`contact` extraction templates.
      operationId: webScrape
      parameters: []
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/WebScrapeDto'
            examples:
              default:
                summary: Default
                value:
                  url: https://example.com
              customOptions:
                summary: Custom options
                value:
                  url: https://example.com
                  format:
                    - markdown
                    - json
                  stealth: true
                  waitTime: 2.5
              withProxy:
                summary: With proxy
                value:
                  url: https://example.com
                  proxyCountry: gb
              withProxyAuto:
                summary: Proxy on demand (auto)
                value:
                  url: https://example.com
                  proxyMode: auto
              withTemplate:
                summary: With extraction template
                value:
                  url: https://testingurl.dev/scraping/ecommerce/product/1
                  extractionMode: template
                  template: product
      responses:
        '200':
          description: Successfully scraped webpage
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/WebScrapeResponseDto'
              examples:
                default:
                  summary: Default scrape
                  value:
                    timestamp: 1778737930991
                    apiStatus: success
                    apiCode: 200
                    meta:
                      url: https://example.com
                      device: desktop
                      format:
                        - html-llm
                      fileOutput: false
                      blockAds: true
                      renderJS: true
                      stealth: false
                      waitTime: 0
                      extractionMode: default
                      proxyMode: 'false'
                      proxyUsed: false
                      test:
                        id: abc123
                    data: |-
                      # Example Domain

                      This domain is for use in illustrative examples...
                templateContact:
                  summary: 'extractionMode: template, template: contact'
                  value:
                    timestamp: 1786985696914
                    apiStatus: success
                    apiCode: 200
                    meta:
                      url: https://testingurl.dev/contact
                      device: desktop
                      format:
                        - json
                      fileOutput: false
                      blockAds: true
                      renderJS: true
                      stealth: false
                      waitTime: 1
                      extractionMode: template
                      template: contact
                      proxyMode: 'false'
                      proxyUsed: false
                      test:
                        id: 8f3a07aa-8b8d-4f21-b0d6-706d3fea5dc0
                    data:
                      json:
                        contact:
                          companyName: TestingURL.dev
                          locations:
                            - id: london
                              label: London
                              address: 221B Baker Street
                              mapUrl: >-
                                https://www.google.com/maps/search/?api=1&query=221B%20Baker%20Street
                              phones:
                                - value: +1-555-0102
                              emails:
                                - value: hello@testingurl.dev
                                  label: general
                                - value: sales@testingurl.dev
                                  label: sales
                          contactChannels:
                            emails:
                              - value: hello@testingurl.dev
                                label: general
                              - value: sales@testingurl.dev
                                label: sales
                            phones:
                              - value: +1-555-0102
                            forms:
                              - https://testingurl.dev/contact
                          socialProfiles:
                            - platform: github
                              url: https://github.com/geekflare/testingurl
                          extractionMeta:
                            fieldsFound:
                              - companyName
                              - locations
                              - address
                              - phones
                              - emails
                              - forms
                              - socialProfiles
                            fieldsMissing:
                              - hours
                              - chatUrl
                              - contactPersons
                        raw: ...
                templateProduct:
                  summary: 'extractionMode: template, template: product'
                  value:
                    timestamp: 1786985781147
                    apiStatus: success
                    apiCode: 200
                    meta:
                      url: https://testingurl.dev/scraping/ecommerce/product/1
                      device: desktop
                      format:
                        - json
                      fileOutput: false
                      blockAds: true
                      renderJS: true
                      stealth: false
                      waitTime: 1
                      extractionMode: template
                      template: product
                      proxyMode: 'false'
                      proxyUsed: false
                      test:
                        id: 22c8fcbf-dcf9-42e5-aaec-a0a3df4b334a
                    data:
                      json:
                        product:
                          title: Vertex Air Laptop 107
                          brand: Vertex
                          category: laptops
                          url: https://testingurl.dev/scraping/ecommerce/product/1
                          description: Designed with a minimalist aesthetic.
                          aggregateRating:
                            ratingValue: 2
                            bestRating: 5
                            reviewCount: 19
                          variants:
                            - sku: TU-1
                              price:
                                amount: 86
                                currency: USD
                              availability:
                                inStock: true
                                text: https://schema.org/InStock
                              images:
                                - url: >-
                                    https://testingurl.dev/assets/placeholder-product.svg
        '400':
          description: Invalid URL.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/BaseErrorResponseDto'
              example:
                timestamp: 1700000000000
                apiStatus: failure
                apiCode: 400
                message: INVALID_URL
                details: The URL must be a valid HTTP or HTTPS URL.
        '422':
          description: Unable to connect to the target website.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/BaseErrorResponseDto'
              example:
                timestamp: 1700000000000
                apiStatus: failure
                apiCode: 422
                message: UNABLE_TO_CONNECT
                details: >-
                  The destination server could not be resolved, refused the
                  connection, timed out, or is redirecting indefinitely.
        '500':
          description: Crawling failed.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/BaseErrorResponseDto'
              example:
                timestamp: 1700000000000
                apiStatus: failure
                apiCode: 500
                message: CRAWL_FAILED
                details: >-
                  Our crawling service encountered an error while attempting to
                  fetch data from the specified URL.
components:
  schemas:
    WebScrapeDto:
      type: object
      properties:
        url:
          type: string
          description: Target URL
          example: https://example.com
        device:
          type: string
          description: Device type to emulate. Defaults to desktop.
          example: desktop
          enum:
            - desktop
            - mobile
          default: desktop
        blockAds:
          type: boolean
          description: Whether to block ads
          example: true
          default: true
        renderJS:
          type: boolean
          description: >-
            Whether to render JavaScript. If omitted, rendering is automatic:
            the page is fetched without a browser first, and JavaScript is only
            rendered if the page needs it. Set explicitly to true or false to
            force rendering on or off.
        proxyMode:
          description: >-
            Controls when a proxy is used. `false` doesn't use a proxy
            (default), `auto` tries without a proxy first and retries through
            one if the site blocks the request, `true` always uses a proxy.
            `proxyMode` is optional — to route through a proxy in a specific
            country, you can set `proxyCountry` on its own.
          default: false
          oneOf:
            - type: boolean
              example: true
            - type: string
              enum:
                - auto
              example: auto
        proxyCountry:
          type: string
          description: >-
            Country code (ISO alpha-2) to route the request through a proxy in.
            Can be used on its own — `proxyMode` is not required. Combine it
            with `proxyMode` if you want control over when the proxy is used.
          example: us
        format:
          type: array
          description: >-
            Format(s) of the scraped result. Comma-separated or array. Defaults
            to html-llm. markdown is recommended for most use cases.
          example: markdown,json
          default:
            - html-llm
          items:
            type: string
            enum:
              - html
              - markdown
              - json
              - markdown-llm
              - html-llm
              - text
              - text-llm
        fileOutput:
          type: boolean
          description: Whether to get response in file format
          example: false
          default: false
        stealth:
          type: boolean
          description: >-
            Enable stealth mode to bypass basic bot detection (removes webdriver
            signals, patches navigator properties)
          example: false
          default: false
        waitTime:
          type: number
          description: >-
            Seconds to wait after page load before capturing content. Helps
            bypass lazy-loaded content and bot checks.
          example: 2.5
          default: 0
        extractionMode:
          type: string
          description: >-
            Extraction mode. JSON output is used automatically when an
            extraction mode is requested — you don't need to set `format` to
            `json`. Set to `template` to use a ready-made extraction template
            instead of a custom schema — see the `template` field.
          example: default
          enum:
            - default
            - cssSchema
            - xpathSchema
            - template
          default: default
        template:
          type: string
          description: >-
            Extraction template to use when extractionMode is `template`
            (ignored otherwise, and has no effect unless extractionMode is set
            to `template`). Accepts `product` (extracts product info: title,
            brand, pricing, availability, images, ratings) or `contact`
            (extracts contact info: company, locations, emails, phones, social
            profiles).
        extractionSchema:
          description: Extraction schema (optional in default mode, required in css/xpath)
          examples:
            default:
              summary: Default Mode Schema
              value:
                name: Quick Fields
                fields:
                  - title: Category
                    value: Electronics
                  - title: Country
                    value: India
            cssSchema:
              summary: CSS Schema
              value:
                name: Product Schema
                baseSelector: .product
                fields:
                  - name: title
                    selector: h1.product-title
                    type: text
                  - name: price
                    selector: .price
                    type: text
            xpathSchema:
              summary: XPath Schema
              value:
                name: Article Schema
                baseSelector: //div[@class='article']
                fields:
                  - name: title
                    selector: //h1/text()
                    type: text
          allOf:
            - $ref: '#/components/schemas/ExtractionSchemaDto'
        aiPrompt:
          description: >-
            Ask AI to extract or analyze the scraped page. Always runs against
            the Markdown of the page regardless of the format field. Adds +7
            credits on top of the base scraping cost.
          oneOf:
            - $ref: '#/components/schemas/PromptAiPromptDto'
            - $ref: '#/components/schemas/SchemaAiPromptDto'
            - $ref: '#/components/schemas/ListingAiPromptDto'
            - $ref: '#/components/schemas/SummaryAiPromptDto'
            - $ref: '#/components/schemas/SentimentAiPromptDto'
            - $ref: '#/components/schemas/KeywordsAiPromptDto'
          discriminator:
            propertyName: type
            mapping:
              prompt: '#/components/schemas/PromptAiPromptDto'
              schema: '#/components/schemas/SchemaAiPromptDto'
              listing: '#/components/schemas/ListingAiPromptDto'
              summary: '#/components/schemas/SummaryAiPromptDto'
              sentiment: '#/components/schemas/SentimentAiPromptDto'
              keywords: '#/components/schemas/KeywordsAiPromptDto'
          examples:
            prompt:
              summary: Open-ended Question
              value:
                type: prompt
                query: What is the return policy?
            schema:
              summary: Custom JSON Schema
              value:
                type: schema
                schema:
                  type: object
                  properties:
                    title:
                      type: string
                    price:
                      type: number
                    currency:
                      type: string
                    inStock:
                      type: boolean
            product:
              summary: Product Extraction
              value:
                type: product
            listing:
              summary: Listing Extraction
              value:
                type: listing
                itemSchema:
                  type: object
                  properties:
                    name:
                      type: string
                    price:
                      type: number
                maxItems: 20
            summary:
              summary: Summary
              value:
                type: summary
                style: bullets
                focus: pricing
            contact:
              summary: Contact Info
              value:
                type: contact
            sentiment:
              summary: Sentiment Analysis
              value:
                type: sentiment
                aspects:
                  - sound quality
                  - battery life
                  - comfort
                  - price
            keywords:
              summary: Keywords & Entities
              value:
                type: keywords
                maxKeywords: 10
                includeEntities: true
      required:
        - url
    WebScrapeResponseDto:
      type: object
      properties:
        timestamp:
          type: number
          description: Timestamp of the request in milliseconds
          example: 1788851167291
        apiStatus:
          type: string
          description: API status message
          example: success
          enum:
            - success
            - failure
        apiCode:
          type: number
          description: API status code
          example: 200
        meta:
          description: Metadata about the request
          allOf:
            - $ref: '#/components/schemas/WebScrapeMetaDto'
        data:
          description: Scraped data (URL or inline content depending on output)
          example: https://example.com/9bulgk075ed9m3vhua5vcrp0.html
          oneOf:
            - type: string
            - type: object
        aiResult:
          type: object
          description: >-
            AI extraction/analysis result. Shape depends on aiPrompt.type.
            Omitted when aiPrompt was not provided.
      required:
        - timestamp
        - apiStatus
        - apiCode
        - meta
        - data
    BaseErrorResponseDto:
      type: object
      properties:
        timestamp:
          type: number
          description: Timestamp of the request in milliseconds
          example: 1778737930991
        apiStatus:
          type: string
          description: API status message
          example: success
          enum:
            - success
            - failure
        apiCode:
          type: number
          description: API status code
          example: 200
        message:
          type: string
          description: Error message
          example: Invalid URL provided
        details:
          type: string
          description: Detailed error information
          example: The URL must be a valid HTTP or HTTPS URL
      required:
        - timestamp
        - apiStatus
        - apiCode
        - message
    ExtractionSchemaDto:
      type: object
      properties:
        name:
          type: string
          description: Name/Label for this extraction schema
          example: Product Schema
        baseSelector:
          type: string
          description: Base selector for scoping extraction (css/xpath only)
          example: .product
        fields:
          type: array
          description: List of fields to extract
          items:
            oneOf:
              - $ref: '#/components/schemas/DefaultExtractionFieldDto'
              - $ref: '#/components/schemas/SelectorExtractionFieldDto'
      required:
        - name
        - fields
    PromptAiPromptDto:
      type: object
      properties:
        type:
          type: string
          description: AI extraction mode
          enum:
            - prompt
            - schema
            - listing
            - summary
            - sentiment
            - keywords
          example: prompt
        query:
          type: string
          description: Open-ended question about the page
          example: What is the return policy?
      required:
        - type
        - query
    SchemaAiPromptDto:
      type: object
      properties:
        type:
          type: string
          description: AI extraction mode
          enum:
            - prompt
            - schema
            - listing
            - summary
            - sentiment
            - keywords
          example: prompt
        schema:
          type: object
          description: JSON Schema-like object describing fields to extract
          example:
            type: object
            properties:
              title:
                type: string
              price:
                type: number
      required:
        - type
        - schema
    ListingAiPromptDto:
      type: object
      properties:
        type:
          type: string
          description: AI extraction mode
          enum:
            - prompt
            - schema
            - listing
            - summary
            - sentiment
            - keywords
          example: prompt
        itemSchema:
          type: object
          description: >-
            JSON Schema-like object describing each item to extract from a
            listing/category page
          example:
            type: object
            properties:
              name:
                type: string
              price:
                type: number
        maxItems:
          type: number
          description: Maximum number of items to extract
          default: 20
          example: 20
      required:
        - type
        - itemSchema
    SummaryAiPromptDto:
      type: object
      properties:
        type:
          type: string
          description: AI extraction mode
          enum:
            - prompt
            - schema
            - listing
            - summary
            - sentiment
            - keywords
          example: prompt
        style:
          type: string
          description: Summary style
          enum:
            - paragraph
            - bullets
            - tldr
          default: paragraph
        focus:
          type: string
          description: Only summarize the parts of the content relevant to this focus area
          example: pricing
        maxLength:
          type: number
          description: Sentence count (paragraph/tldr) or bullet count (bullets)
          default: 5
          example: 5
      required:
        - type
    SentimentAiPromptDto:
      type: object
      properties:
        type:
          type: string
          description: AI extraction mode
          enum:
            - prompt
            - schema
            - listing
            - summary
            - sentiment
            - keywords
          example: prompt
        aspects:
          description: >-
            Aspects to score individually (aspect-based sentiment). If omitted,
            only overall sentiment is returned.
          example:
            - sound quality
            - battery life
            - comfort
            - price
          type: array
          items:
            type: string
      required:
        - type
    KeywordsAiPromptDto:
      type: object
      properties:
        type:
          type: string
          description: AI extraction mode
          enum:
            - prompt
            - schema
            - listing
            - summary
            - sentiment
            - keywords
          example: prompt
        maxKeywords:
          type: number
          description: Maximum number of keywords to return
          default: 10
          example: 10
        includeEntities:
          type: boolean
          description: >-
            Also extract named entities (organizations, people, dates, laws,
            locations)
          default: false
      required:
        - type
    WebScrapeMetaDto:
      type: object
      properties:
        url:
          type: string
          description: The target URL that was scraped
          example: https://example.com
        device:
          type: string
          description: Device type used
          example: desktop
          enum:
            - desktop
            - mobile
        format:
          type: array
          description: Output format(s) of the result
          example:
            - html-llm
          items:
            type: string
            enum:
              - html
              - markdown
              - json
              - markdown-llm
              - html-llm
              - text
              - text-llm
        fileOutput:
          type: boolean
          description: Whether to get response in file format
          example: false
        blockAds:
          type: boolean
          description: Whether ads were blocked
          example: true
        renderJS:
          type: boolean
          description: >-
            Whether JavaScript was rendered for this request (resolved
            automatically unless explicitly set)
          example: true
        stealth:
          type: boolean
          description: Whether stealth mode was enabled
          example: false
        proxyMode:
          type: string
          description: >-
            Proxy mode requested for this request, echoed as a string ("false",
            "auto", or "true")
          example: 'false'
        proxyUsed:
          type: boolean
          description: >-
            Whether a proxy was actually used for this request. When proxyMode
            is `auto`, this depends on whether the site blocked the initial
            request.
          example: false
        waitTime:
          type: number
          description: >-
            Seconds to wait after page load before capturing content. Helps
            bypass lazy-loaded content and bot checks.
          example: 2.5
          default: 0
        proxyCountry:
          type: string
          description: Proxy country used, if any
        extractionMode:
          type: string
          description: >-
            Extraction mode used for this request. When an extraction mode is
            requested, the result is returned as JSON.
          example: default
        template:
          type: string
          description: Extraction template used, if extractionMode was `template`
          example: product
          enum:
            - product
            - contact
        extractionSchema:
          description: Extraction schema (optional in default mode, required in css/xpath)
          examples:
            default:
              summary: Default Mode Schema
              value:
                name: Quick Fields
                fields:
                  - title: Category
                    value: Electronics
                  - title: Country
                    value: India
          allOf:
            - $ref: '#/components/schemas/ExtractionSchemaDto'
        test:
          description: Test details object
          allOf:
            - $ref: '#/components/schemas/TestMetaDto'
        aiPromptType:
          type: string
          description: The aiPrompt.type used for this request, if any
          example: prompt
      required:
        - url
        - device
        - format
        - fileOutput
        - blockAds
        - renderJS
        - stealth
        - proxyMode
        - proxyUsed
        - waitTime
        - extractionMode
        - extractionSchema
        - test
    DefaultExtractionFieldDto:
      type: object
      properties:
        title:
          type: string
          description: Title/key of the extracted field
          example: Product Name
        value:
          type: object
          description: Static value to assign to this field
          example: Some hardcoded string
      required:
        - title
        - value
    SelectorExtractionFieldDto:
      type: object
      properties:
        name:
          type: string
          description: Field name in the extracted JSON
          example: title
        selector:
          type: string
          description: Selector or XPath to extract value
          example: h1.product-title
        type:
          type: string
          description: Type of data to extract
          example: text
        attribute:
          type: string
          description: If type=attr, specify attribute name
          example: href
        fields:
          description: Nested fields
          type: array
          items:
            $ref: '#/components/schemas/SelectorExtractionFieldDto'
      required:
        - name
        - selector
        - type
    TestMetaDto:
      type: object
      properties:
        id:
          type: string
          description: Unique test identifier
          example: mxqx9v9y0742lap6altwdteqd28t23nq
      required:
        - id
  securitySchemes:
    x-api-key:
      type: apiKey
      in: header
      name: x-api-key
      description: API Key required for all endpoints

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.