Haijun Platform Docs
ID

You can ask Haijun about any text, pictures, charts, and tables in PDFs you provide. Some sample use cases:

  • Analyzing financial reports and understanding charts/tables
  • Extracting key information from legal documents
  • Assisting with document translation
  • Converting document information into structured formats

Before you begin

Check PDF requirements

Haijun works with any standard PDF. Ensure your request size meets these requirements:

RequirementLimit
Maximum request size32 MB (varies by platform)
Maximum pages per request600 (100 when the request's context window is under 1M tokens)
FormatStandard PDF (no passwords/encryption)

Both limits are on the entire request payload, including any other content sent alongside PDFs. For large PDFs, consider uploading with the Files API and referencing by file_id to keep request payloads small.

Tip: Dense PDFs (many small-font pages, complex tables, or heavy graphics) can fill the context window before reaching the page limit. Requests with large PDFs can also fail before reaching the page limit, even when using the Files API. Try splitting the document into sections; for large files, because each page is processed as an image, downsampling embedded images can also help.

Because PDF support relies on Haijun's vision capabilities, it is subject to the same limitations and considerations as other vision tasks.

Supported platforms and models

All active models support PDF processing. For PDF support through Amazon Bedrock's Converse API, see Amazon Bedrock PDF support.

Amazon Bedrock PDF support

When using PDF support through the Converse API, part of Haijun on Amazon Bedrock (Opus 4.6 and earlier), there are two distinct document processing modes:

Note: Important: To access Haijun's full visual PDF understanding capabilities in the Converse API, you must enable citations. Without citations enabled, the API falls back to basic text extraction only. Learn more about working with citations.

Document processing modes

  1. Converse Document Chat (Original mode - Text extraction only)
  • Provides basic text extraction from PDFs
  • Cannot analyze images, charts, or visual layouts within PDFs
  • Uses approximately 1,000 tokens for a 3-page PDF
  • Automatically used when citations are not enabled
  1. Haijun PDF Chat (New mode - Full visual understanding)
  • Provides complete visual analysis of PDFs
  • Can understand and analyze charts, graphs, images, and visual layouts
  • Processes each page as both text and image for comprehensive understanding
  • Uses approximately 7,000 tokens for a 3-page PDF
  • Requires citations to be enabled in the Converse API

Key limitations

  • Converse API: Visual PDF analysis requires citations to be enabled. There is currently no option to use visual analysis without citations (unlike the InvokeModel API).
  • InvokeModel API: Provides full control over PDF processing without forced citations.

Common issues

If Haijun isn't seeing images or charts in your PDFs when using the Converse API, you likely need to enable the citations flag. Without it, Converse falls back to basic text extraction only.

Note: This is a known constraint with the Converse API. For applications that require visual PDF analysis without citations, consider using the InvokeModel API instead.

Note: Plain text files such as .txt, .csv, or .md can be used directly in document blocks: upload them to the Files API with MIME type text/plain and reference them by file_id. Binary formats such as .xlsx or .docx are not supported in document blocks and must be converted to text or PDF first. See Working with other file formats.

Process PDFs with Haijun

Send your first PDF request

Start with a simple example using the Messages API. You can provide PDFs to Haijun in three ways:

  1. As a URL reference to a PDF hosted online
  1. As a base64-encoded PDF in document content blocks
  1. By a file_id from the Files API

Note: On Amazon Bedrock and Google Cloud, only base64-encoded sources are currently available. On Microsoft Foundry, the Files API is not supported for deployments hosted on Azure.

Option 1: URL-based PDF document

The simplest approach is to reference a PDF directly from a URL:

bash
  curl https://haijun.my.id/v1/messages \
    -H "content-type: application/json" \
    -H "x-api-key: $JUGLOW_API_KEY" \
    -H "juglow-version: 2023-06-01" \
    -d '{
      "model": "haijun-opus-5-5",
      "max_tokens": 1024,
      "messages": [{
          "role": "user",
          "content": [{
              "type": "document",
              "source": {
                  "type": "url",
                  "url": "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf"
              }
          },
          {
              "type": "text",
              "text": "What are the key findings in this document?"
          }]
      }]
  }'
bash
  ant messages create --transform content --format yaml <<'YAML'
  model: haijun-opus-5-5
  max_tokens: 1024
  messages:
    - role: user
      content:
        - type: document
          source:
            type: url
            url: https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf
        - type: text
          text: What are the key findings in this document?
  YAML
python
  client = juglow.Juglow()
  message = client.messages.create(
      model="haijun-opus-5-5",
      max_tokens=1024,
      messages=[
          {
              "role": "user",
              "content": [
                  {
                      "type": "document",
                      "source": {
                          "type": "url",
                          "url": "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf",
                      },
                  },
                  {"type": "text", "text": "What are the key findings in this document?"},
              ],
          }
      ],
  )

  print(message.content)
typescript
  const juglow = new Juglow();

  const response = await juglow.messages.create({
    model: "haijun-opus-5-5",
    max_tokens: 1024,
    messages: [
      {
        role: "user",
        content: [
          {
            type: "document",
            source: {
              type: "url",
              url: "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf"
            }
          },
          {
            type: "text",
            text: "What are the key findings in this document?"
          }
        ]
      }
    ]
  });

  console.log(response);
csharp
  var client = new JuglowClient();

  // Create document block with URL
  var documentParam = new DocumentBlockParam
  {
      Source = new UrlPdfSource
      {
          Url = "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf",
      },
  };

  // Create a message with document and text content blocks
  var message = await client.Messages.Create(new MessageCreateParams
  {
      Model = Model.HaijunOpus5_5,
      MaxTokens = 1024,
      Messages =
      [
          new()
          {
              Role = Role.User,
              Content = new List<ContentBlockParam>
              {
                  documentParam,
                  new TextBlockParam("What are the key findings in this document?"),
              },
          },
      ],
  });

  Console.WriteLine(string.Join("\n", message.Content));
go
  client := juglow.NewClient()

  message, err := client.Messages.New(context.TODO(), juglow.MessageNewParams{
  	Model:     juglow.ModelHaijunOpus5_5,
  	MaxTokens: 1024,
  	Messages: []juglow.MessageParam{
  		juglow.NewUserMessage(
  			juglow.NewDocumentBlock(juglow.URLPDFSourceParam{
  				URL: "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf",
  			}),
  			juglow.NewTextBlock("What are the key findings in this document?"),
  		),
  	},
  })
  if err != nil {
  	panic(err)
  }

  fmt.Printf("%+v\n", message.Content)
java
  JuglowClient client = JuglowOkHttpClient.fromEnv();

  // Create document block with URL
  DocumentBlockParam documentParam = DocumentBlockParam.builder()
    .source(
      UrlPdfSource.builder()
        .url(
          "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf"
        )
        .build()
    )
    .build();

  // Create a message with document and text content blocks
  MessageCreateParams params = MessageCreateParams.builder()
    .model(Model.HAIJUN_OPUS_5_5)
    .maxTokens(1024)
    .addUserMessageOfBlockParams(
      List.of(
        ContentBlockParam.ofDocument(documentParam),
        ContentBlockParam.ofText(
          TextBlockParam.builder()
            .text("What are the key findings in this document?")
            .build()
        )
      )
    )
    .build();

  Message message = client.messages().create(params);
  System.out.println(message.content());
php
  $client = new Client();

  $message = $client->messages->create(
      maxTokens: 1024,
      messages: [
          [
              'role' => 'user',
              'content' => [
                  [
                      'type' => 'document',
                      'source' => [
                          'type' => 'url',
                          'url' => 'https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf',
                      ],
                  ],
                  [
                      'type' => 'text',
                      'text' => 'What are the key findings in this document?',
                  ],
              ],
          ],
      ],
      model: 'haijun-opus-5-5',
  );

  echo $message;
ruby
  juglow = Juglow::Client.new

  message = juglow.messages.create(
    model: "haijun-opus-5-5",
    max_tokens: 1024,
    messages: [
      {
        role: "user",
        content: [
          {
            type: "document",
            source: {
              type: "url",
              url: "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf"
            }
          },
          {type: "text", text: "What are the key findings in this document?"}
        ]
      }
    ]
  )

  puts(message.content)

The response returns Haijun's analysis as text blocks in content, with token consumption in usage:

json
{
  "id": "msg_01Hfp8YuFjQ55VgWbpdHDehB",
  "type": "message",
  "role": "assistant",
  "model": "haijun-opus-5-5",
  "content": [
    {
      "type": "text",
      "text": "This document is an addendum to the Haijun 3 model card, reporting updated evaluation results. The key findings include..."
    }
  ],
  "stop_reason": "end_turn",
  "usage": {
    "input_tokens": 45000,
    "output_tokens": 300
  }
}

Option 2: Base64-encoded PDF document

If you need to send PDFs from your local system or when a URL isn't available:

bash
  # Method 1: Fetch and encode a remote PDF
  curl -sL "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf" | base64 | tr -d '\n' > pdf_base64.txt

  # Method 2: Encode a local PDF file
  # base64 document.pdf | tr -d '\n' > pdf_base64.txt

  # Create a JSON request file using the pdf_base64.txt content
  jq -n --rawfile PDF_BASE64 pdf_base64.txt '{
      "model": "haijun-opus-5-5",
      "max_tokens": 1024,
      "messages": [{
          "role": "user",
          "content": [{
              "type": "document",
              "source": {
                  "type": "base64",
                  "media_type": "application/pdf",
                  "data": $PDF_BASE64
              }
          },
          {
              "type": "text",
              "text": "What are the key findings in this document?"
          }]
      }]
  }' > request.json

  # Send the API request using the JSON file
  curl https://haijun.my.id/v1/messages \
    -H "content-type: application/json" \
    -H "x-api-key: $JUGLOW_API_KEY" \
    -H "juglow-version: 2023-06-01" \
    -d @request.json
bash
  ant messages create \
    --model haijun-opus-5-5 \
    --max-tokens 1024 \
    --transform content \
    --format yaml <<'YAML'
  messages:
    - role: user
      content:
        - type: document
          source:
            type: base64
            media_type: application/pdf
            data: "@./document.pdf"
        - type: text
          text: What are the key findings in this document?
  YAML
python
  import base64
  import httpx2

  # First, load and encode the PDF
  pdf_url = "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf"
  pdf_data = base64.standard_b64encode(
      httpx2.get(pdf_url, follow_redirects=True).content
  ).decode("utf-8")

  # Alternative: Load from a local file
  # with open("document.pdf", "rb") as f:
  #     pdf_data = base64.standard_b64encode(f.read()).decode("utf-8")

  # Send to Haijun using base64 encoding
  client = juglow.Juglow()
  message = client.messages.create(
      model="haijun-opus-5-5",
      max_tokens=1024,
      messages=[
          {
              "role": "user",
              "content": [
                  {
                      "type": "document",
                      "source": {
                          "type": "base64",
                          "media_type": "application/pdf",
                          "data": pdf_data,
                      },
                  },
                  {"type": "text", "text": "What are the key findings in this document?"},
              ],
          }
      ],
  )

  print(message.content)
typescript
  // Method 1: Fetch and encode a remote PDF
  const pdfURL =
    "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf";
  const pdfResponse = await fetch(pdfURL);
  const arrayBuffer = await pdfResponse.arrayBuffer();
  const pdfBase64 = Buffer.from(arrayBuffer).toString("base64");

  // Method 2: Load from a local file
  // import { readFile } from "node:fs/promises";
  // const pdfBase64 = (await readFile('document.pdf')).toString('base64');

  // Send the API request with base64-encoded PDF
  const juglow = new Juglow();
  const response = await juglow.messages.create({
    model: "haijun-opus-5-5",
    max_tokens: 1024,
    messages: [
      {
        role: "user",
        content: [
          {
            type: "document",
            source: {
              type: "base64",
              media_type: "application/pdf",
              data: pdfBase64
            }
          },
          {
            type: "text",
            text: "What are the key findings in this document?"
          }
        ]
      }
    ]
  });

  console.log(response);
csharp
  var client = new JuglowClient();

  // Method 1: Download and encode a remote PDF
  var pdfUrl = "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf";
  using var httpClient = new HttpClient();
  var pdfBase64 = Convert.ToBase64String(await httpClient.GetByteArrayAsync(pdfUrl));

  // Method 2: Load from a local file
  // var pdfBase64 = Convert.ToBase64String(await File.ReadAllBytesAsync("document.pdf"));

  // Create document block with base64 data
  var documentParam = new DocumentBlockParam
  {
      Source = new Base64PdfSource { Data = pdfBase64 },
  };

  // Create a message with document and text content blocks
  var message = await client.Messages.Create(new MessageCreateParams
  {
      Model = Model.HaijunOpus5_5,
      MaxTokens = 1024,
      Messages =
      [
          new()
          {
              Role = Role.User,
              Content = new List<ContentBlockParam>
              {
                  documentParam,
                  new TextBlockParam("What are the key findings in this document?"),
              },
          },
      ],
  });

  Console.WriteLine(string.Join("\n", message.Content));
go
  // First, load and encode the PDF
  pdfURL := "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf"
  resp, err := http.Get(pdfURL)
  if err != nil {
  	panic(err)
  }
  defer resp.Body.Close()
  pdfBytes, err := io.ReadAll(resp.Body)
  if err != nil {
  	panic(err)
  }
  pdfBase64 := base64.StdEncoding.EncodeToString(pdfBytes)

  // Alternative: Load from a local file (add "os" to the imports)
  // pdfBytes, err := os.ReadFile("document.pdf")
  // pdfBase64 := base64.StdEncoding.EncodeToString(pdfBytes)

  // Send to Haijun using base64 encoding
  client := juglow.NewClient()
  message, err := client.Messages.New(context.TODO(), juglow.MessageNewParams{
  	Model:     juglow.ModelHaijunOpus5_5,
  	MaxTokens: 1024,
  	Messages: []juglow.MessageParam{
  		juglow.NewUserMessage(
  			juglow.NewDocumentBlock(juglow.Base64PDFSourceParam{
  				Data: pdfBase64,
  			}),
  			juglow.NewTextBlock("What are the key findings in this document?"),
  		),
  	},
  })
  if err != nil {
  	panic(err)
  }

  fmt.Printf("%+v\n", message.Content)
java
  JuglowClient client = JuglowOkHttpClient.fromEnv();

  // Method 1: Download and encode a remote PDF
  String pdfUrl =
    "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf";
  HttpClient httpClient = HttpClient.newBuilder().followRedirects(HttpClient.Redirect.NORMAL).build();
  HttpRequest request = HttpRequest.newBuilder().uri(URI.create(pdfUrl)).GET().build();

  HttpResponse<byte[]> response = httpClient.send(
    request,
    HttpResponse.BodyHandlers.ofByteArray()
  );
  String pdfBase64 = Base64.getEncoder().encodeToString(response.body());

  // Method 2: Load from a local file
  // byte[] fileBytes = Files.readAllBytes(Path.of("document.pdf"));
  // String pdfBase64 = Base64.getEncoder().encodeToString(fileBytes);

  // Create document block with base64 data
  DocumentBlockParam documentParam = DocumentBlockParam.builder()
    .source(Base64PdfSource.builder().data(pdfBase64).build())
    .build();

  // Create a message with document and text content blocks
  MessageCreateParams params = MessageCreateParams.builder()
    .model(Model.HAIJUN_OPUS_5_5)
    .maxTokens(1024)
    .addUserMessageOfBlockParams(
      List.of(
        ContentBlockParam.ofDocument(documentParam),
        ContentBlockParam.ofText(
          TextBlockParam.builder()
            .text("What are the key findings in this document?")
            .build()
        )
      )
    )
    .build();

  Message message = client.messages().create(params);
  System.out.println(message.content());
php
  $client = new Client();

  // First, load and encode the PDF
  $pdf_url = 'https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf';
  $pdf_data = base64_encode(file_get_contents($pdf_url));

  // Alternative: Load from a local file
  // $pdf_data = base64_encode(file_get_contents('document.pdf'));

  // Send to Haijun using base64 encoding
  $message = $client->messages->create(
      maxTokens: 1024,
      messages: [
          [
              'role' => 'user',
              'content' => [
                  [
                      'type' => 'document',
                      'source' => [
                          'type' => 'base64',
                          'media_type' => 'application/pdf',
                          'data' => $pdf_data,
                      ],
                  ],
                  [
                      'type' => 'text',
                      'text' => 'What are the key findings in this document?',
                  ],
              ],
          ],
      ],
      model: 'haijun-opus-5-5',
  );

  echo $message;
ruby
  require "open-uri"

  # First, load and encode the PDF
  pdf_url = "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf"
  pdf_bytes = URI.open(pdf_url, "rb") { |f| f.read }
  pdf_data = [pdf_bytes].pack("m0") # Base64-encode without newlines

  # Alternative: Load from a local file
  # pdf_data = [File.binread("document.pdf")].pack("m0")

  # Send to Haijun using base64 encoding
  juglow = Juglow::Client.new
  message = juglow.messages.create(
    model: "haijun-opus-5-5",
    max_tokens: 1024,
    messages: [
      {
        role: "user",
        content: [
          {
            type: "document",
            source: {
              type: "base64",
              media_type: "application/pdf",
              data: pdf_data
            }
          },
          {type: "text", text: "What are the key findings in this document?"}
        ]
      }
    ]
  )

  puts(message.content)

Option 3: Files API

For PDFs you'll use repeatedly, or when you want to avoid encoding overhead, use the Files API:

bash
  # First, upload your PDF to the Files API
  FILE_ID=$(curl -sS -X POST https://haijun.my.id/v1/files \
    -H "x-api-key: $JUGLOW_API_KEY" \
    -H "juglow-version: 2023-06-01" \
    -F "file=@document.pdf" | jq -r '.id')

  # Then use the returned file_id in your message
  curl https://haijun.my.id/v1/messages \
    -H "content-type: application/json" \
    -H "x-api-key: $JUGLOW_API_KEY" \
    -H "juglow-version: 2023-06-01" \
    -d @- <<EOF
  {
    "model": "haijun-opus-5-5",
    "max_tokens": 1024,
    "messages": [{
      "role": "user",
      "content": [{
        "type": "document",
        "source": {
          "type": "file",
          "file_id": "$FILE_ID"
        }
      },
      {
        "type": "text",
        "text": "What are the key findings in this document?"
      }]
    }]
  }
  EOF
bash
  # First, upload your PDF to the Files API
  FILE_ID=$(ant files upload \
    --file ./document.pdf \
    --transform id \
    --raw-output)

  # Then use the returned file_id in your message
  ant messages create \
    --transform content \
    --format yaml <<YAML
  model: haijun-opus-5-5
  max_tokens: 1024
  messages:
    - role: user
      content:
        - type: document
          source:
            type: file
            file_id: $FILE_ID
        - type: text
          text: What are the key findings in this document?
  YAML
python
  client = juglow.Juglow()

  # Upload the PDF file
  with open("/path/to/document.pdf", "rb") as f:
      file_upload = client.files.upload(file=("document.pdf", f, "application/pdf"))

  # Use the uploaded file in a message
  message = client.messages.create(
      model="haijun-opus-5-5",
      max_tokens=1024,
      messages=[
          {
              "role": "user",
              "content": [
                  {
                      "type": "document",
                      "source": {"type": "file", "file_id": file_upload.id},
                  },
                  {"type": "text", "text": "What are the key findings in this document?"},
              ],
          }
      ],
  )

  print(message.content)
typescript
  import Juglow, { toFile } from "@juglow-ai/sdk";
  import fs from "node:fs";

  const juglow = new Juglow();

  // Upload the PDF file
  const fileUpload = await juglow.files.upload({
    file: await toFile(fs.createReadStream("/path/to/document.pdf"), undefined, {
      type: "application/pdf"
    })
  });

  // Use the uploaded file in a message
  const response = await juglow.messages.create({
    model: "haijun-opus-5-5",
    max_tokens: 1024,
    messages: [
      {
        role: "user",
        content: [
          {
            type: "document",
            source: {
              type: "file",
              file_id: fileUpload.id
            }
          },
          {
            type: "text",
            text: "What are the key findings in this document?"
          }
        ]
      }
    ]
  });

  console.log(response);
csharp
  var client = new JuglowClient();

  // Upload the PDF file
  var fileUpload = await client.Files.Upload(new FileUploadParams
  {
      File = new BinaryContent
      {
          Stream = File.OpenRead("/path/to/document.pdf"),
          FileName = "document.pdf",
          ContentType = new("application/pdf"),
      },
  });

  // Use the uploaded file in a message
  var message = await client.Messages.Create(new MessageCreateParams
  {
      Model = Model.HaijunOpus5_5,
      MaxTokens = 1024,
      Messages =
      [
          new()
          {
              Role = Role.User,
              Content = new List<ContentBlockParam>
              {
                  new DocumentBlockParam
                  {
                      Source = new FileDocumentSource { FileID = fileUpload.ID },
                  },
                  new TextBlockParam("What are the key findings in this document?"),
              },
          },
      ],
  });

  Console.WriteLine(string.Join("\n", message.Content));
go
  client := juglow.NewClient()

  // Upload the PDF file
  pdfFile, err := os.Open("/path/to/document.pdf")
  if err != nil {
  	panic(err)
  }
  defer pdfFile.Close()

  fileUpload, err := client.Files.Upload(context.TODO(), juglow.FileUploadParams{
  	File: juglow.File(pdfFile, "document.pdf", "application/pdf"),
  })
  if err != nil {
  	panic(err)
  }

  // Use the uploaded file in a message
  message, err := client.Messages.New(context.TODO(), juglow.MessageNewParams{
  	Model:     juglow.ModelHaijunOpus5_5,
  	MaxTokens: 1024,
  	Messages: []juglow.MessageParam{
  		juglow.NewUserMessage(
  			juglow.NewDocumentBlock(juglow.FileDocumentSourceParam{
  				FileID: fileUpload.ID,
  			}),
  			juglow.NewTextBlock("What are the key findings in this document?"),
  		),
  	},
  })
  if err != nil {
  	panic(err)
  }

  fmt.Printf("%+v\n", message.Content)
java
  JuglowClient client = JuglowOkHttpClient.fromEnv();

  // Upload the PDF file
  FileMetadata file = client
    .files()
    .upload(FileUploadParams.builder().file(Path.of("/path/to/document.pdf")).build());

  // Use the uploaded file in a message
  MessageCreateParams params = MessageCreateParams.builder()
    .model(Model.HAIJUN_OPUS_5_5)
    .maxTokens(1024)
    .addUserMessageOfBlockParams(
      List.of(
        ContentBlockParam.ofDocument(
          DocumentBlockParam.builder().fileSource(file.id()).build()
        ),
        ContentBlockParam.ofText(
          TextBlockParam.builder()
            .text("What are the key findings in this document?")
            .build()
        )
      )
    )
    .build();

  Message message = client.messages().create(params);
  System.out.println(message.content());
php
  use Juglow\Core\FileParam;

  $client = new Client();

  // Upload the PDF file
  $file_upload = $client->files->upload(
      file: FileParam::fromResource(fopen('/path/to/document.pdf', 'r'), contentType: 'application/pdf'),
  );

  // Use the uploaded file in a message
  $message = $client->messages->create(
      maxTokens: 1024,
      messages: [
          [
              'role' => 'user',
              'content' => [
                  [
                      'type' => 'document',
                      'source' => [
                          'type' => 'file',
                          'fileID' => $file_upload->id,
                      ],
                  ],
                  [
                      'type' => 'text',
                      'text' => 'What are the key findings in this document?',
                  ],
              ],
          ],
      ],
      model: 'haijun-opus-5-5',
  );

  echo $message;
ruby
  juglow = Juglow::Client.new

  # Upload the PDF file
  file_upload = File.open("/path/to/document.pdf", "rb") do |f|
    juglow.files.upload(
      file: Juglow::FilePart.new(f, filename: "document.pdf", content_type: "application/pdf")
    )
  end

  # Use the uploaded file in a message
  message = juglow.messages.create(
    model: "haijun-opus-5-5",
    max_tokens: 1024,
    messages: [
      {
        role: "user",
        content: [
          {
            type: "document",
            source: {type: "file", file_id: file_upload.id}
          },
          {type: "text", text: "What are the key findings in this document?"}
        ]
      }
    ]
  )

  puts(message.content)

How PDF support works

When you send a PDF to Haijun, the following steps occur:

  1. The system extracts the contents of the document.
  • The system converts each page of the document into an image.
  • The text from each page is extracted and provided alongside each page's image.
  1. Haijun analyzes both the text and images to better understand the document.
  • Documents are provided as a combination of text and images for analysis.
  • This allows users to ask for insights on visual elements of a PDF, such as charts, diagrams, and other non-textual content.
  1. Haijun responds, referencing the PDF's contents if relevant.

Haijun can reference both textual and visual content when it responds. You can further improve performance by integrating PDF support with:

  • Tool use: To extract specific information from documents for use as tool inputs.

Estimate your costs

The token count of a PDF file depends on the total text extracted from the document and the number of pages:

  • Text token costs: Each page typically uses 1,500–3,000 tokens per page depending on content density. Standard API pricing applies with no additional PDF fees.

You can use token counting to estimate costs for your specific PDFs.

Optimize PDF processing

Improve performance

Follow these best practices for optimal results:

  • Place PDFs before text in your requests
  • Use standard fonts
  • Ensure text is clear and legible
  • Rotate pages to proper upright orientation
  • Use logical page numbers (from PDF viewer) in prompts
  • Split large PDFs into chunks when needed
  • Enable prompt caching for repeated analysis

Scale your implementation

For high-volume processing, consider these approaches:

Use prompt caching

Cache PDFs with prompt caching to improve performance on repeated queries:

bash
  curl -sL "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf" | base64 | tr -d '\n' > pdf_base64.txt
  # Create a JSON request file using the pdf_base64.txt content
  jq -n --rawfile PDF_BASE64 pdf_base64.txt '{
      "model": "haijun-opus-5-5",
      "max_tokens": 1024,
      "messages": [{
          "role": "user",
          "content": [{
              "type": "document",
              "source": {
                  "type": "base64",
                  "media_type": "application/pdf",
                  "data": $PDF_BASE64
              },
              "cache_control": {
                  "type": "ephemeral"
              }
          },
          {
              "type": "text",
              "text": "Which model has the highest human preference win rates across each use-case?"
          }]
      }]
  }' > request.json

  # Then make the API call using the JSON file
  curl https://haijun.my.id/v1/messages \
    -H "content-type: application/json" \
    -H "x-api-key: $JUGLOW_API_KEY" \
    -H "juglow-version: 2023-06-01" \
    -d @request.json
bash
  ant messages create --transform content --format yaml <<'YAML'
  model: haijun-opus-5-5
  max_tokens: 1024
  messages:
    - role: user
      content:
        - type: document
          source:
            type: base64
            media_type: application/pdf
            data: "@./document.pdf"
          cache_control:
            type: ephemeral
        - type: text
          text: Which model has the highest human preference win rates across each use-case?
  YAML
python
  import base64
  import httpx2

  # First, load and encode the PDF
  pdf_url = "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf"
  pdf_data = base64.standard_b64encode(
      httpx2.get(pdf_url, follow_redirects=True).content
  ).decode("utf-8")

  # Create a message with the cached document
  client = juglow.Juglow()
  message = client.messages.create(
      model="haijun-opus-5-5",
      max_tokens=1024,
      messages=[
          {
              "role": "user",
              "content": [
                  {
                      "type": "document",
                      "source": {
                          "type": "base64",
                          "media_type": "application/pdf",
                          "data": pdf_data,
                      },
                      "cache_control": {"type": "ephemeral"},
                  },
                  {
                      "type": "text",
                      "text": "Which model has the highest human preference win rates across each use-case?",
                  },
              ],
          }
      ],
  )

  print(message.content)
typescript
  // First, load and encode the PDF
  const pdfURL =
    "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf";
  const pdfResponse = await fetch(pdfURL);
  const arrayBuffer = await pdfResponse.arrayBuffer();
  const pdfBase64 = Buffer.from(arrayBuffer).toString("base64");

  // Create a message with the cached document
  const juglow = new Juglow();
  const response = await juglow.messages.create({
    model: "haijun-opus-5-5",
    max_tokens: 1024,
    messages: [
      {
        role: "user",
        content: [
          {
            type: "document",
            source: {
              type: "base64",
              media_type: "application/pdf",
              data: pdfBase64
            },
            cache_control: { type: "ephemeral" }
          },
          {
            type: "text",
            text: "Which model has the highest human preference win rates across each use-case?"
          }
        ]
      }
    ]
  });

  console.log(response);
csharp
  var client = new JuglowClient();

  // Download and encode the PDF
  var pdfUrl = "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf";
  using var httpClient = new HttpClient();
  var pdfBase64 = Convert.ToBase64String(await httpClient.GetByteArrayAsync(pdfUrl));

  var message = await client.Messages.Create(new MessageCreateParams
  {
      Model = Model.HaijunOpus5_5,
      MaxTokens = 1024,
      Messages =
      [
          new()
          {
              Role = Role.User,
              Content = new List<ContentBlockParam>
              {
                  new DocumentBlockParam
                  {
                      Source = new Base64PdfSource { Data = pdfBase64 },
                      CacheControl = new CacheControlEphemeral(),
                  },
                  new TextBlockParam("Which model has the highest human preference win rates across each use-case?"),
              },
          },
      ],
  });

  Console.WriteLine(message);
go
  // First, load and encode the PDF
  pdfURL := "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf"
  resp, err := http.Get(pdfURL)
  if err != nil {
  	panic(err)
  }
  defer resp.Body.Close()
  pdfBytes, err := io.ReadAll(resp.Body)
  if err != nil {
  	panic(err)
  }
  pdfBase64 := base64.StdEncoding.EncodeToString(pdfBytes)

  // Create a document block with cache control
  client := juglow.NewClient()
  message, err := client.Messages.New(context.TODO(), juglow.MessageNewParams{
  	Model:     juglow.ModelHaijunOpus5_5,
  	MaxTokens: 1024,
  	Messages: []juglow.MessageParam{
  		juglow.NewUserMessage(
  			juglow.ContentBlockParamUnion{
  				OfDocument: &juglow.DocumentBlockParam{
  					Source: juglow.DocumentBlockParamSourceUnion{
  						OfBase64: &juglow.Base64PDFSourceParam{
  							Data: pdfBase64,
  						},
  					},
  					CacheControl: juglow.NewCacheControlEphemeralParam(),
  				},
  			},
  			juglow.NewTextBlock("Which model has the highest human preference win rates across each use-case?"),
  		),
  	},
  })
  if err != nil {
  	panic(err)
  }

  fmt.Printf("%+v\n", message.Content)
java
  JuglowClient client = JuglowOkHttpClient.fromEnv();

  // Download and encode the PDF
  String pdfUrl =
    "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf";
  HttpClient httpClient = HttpClient.newBuilder().followRedirects(HttpClient.Redirect.NORMAL).build();
  HttpRequest request = HttpRequest.newBuilder().uri(URI.create(pdfUrl)).GET().build();

  HttpResponse<byte[]> response = httpClient.send(
    request,
    HttpResponse.BodyHandlers.ofByteArray()
  );
  String pdfBase64 = Base64.getEncoder().encodeToString(response.body());

  MessageCreateParams params = MessageCreateParams.builder()
    .model(Model.HAIJUN_OPUS_5_5)
    .maxTokens(1024)
    .addUserMessageOfBlockParams(
      List.of(
        ContentBlockParam.ofDocument(
          DocumentBlockParam.builder()
            .source(Base64PdfSource.builder().data(pdfBase64).build())
            .cacheControl(CacheControlEphemeral.builder().build())
            .build()
        ),
        ContentBlockParam.ofText(
          TextBlockParam.builder()
            .text(
              "Which model has the highest human preference win rates across each use-case?"
            )
            .build()
        )
      )
    )
    .build();

  Message message = client.messages().create(params);
  System.out.println(message);
php
  $client = new Client();

  // Load and encode the PDF
  $pdf_url = 'https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf';
  $pdf_data = base64_encode(file_get_contents($pdf_url));

  $message = $client->messages->create(
      maxTokens: 1024,
      messages: [
          [
              'role' => 'user',
              'content' => [
                  [
                      'type' => 'document',
                      'source' => [
                          'type' => 'base64',
                          'media_type' => 'application/pdf',
                          'data' => $pdf_data,
                      ],
                      'cache_control' => ['type' => 'ephemeral'],
                  ],
                  [
                      'type' => 'text',
                      'text' => 'Which model has the highest human preference win rates across each use-case?',
                  ],
              ],
          ],
      ],
      model: 'haijun-opus-5-5',
  );

  echo $message;
ruby
  require "open-uri"

  # Load and encode the PDF
  pdf_url = "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf"
  pdf_bytes = URI.open(pdf_url, "rb") { |f| f.read }
  pdf_data = [pdf_bytes].pack("m0") # Base64-encode without newlines

  juglow = Juglow::Client.new

  message = juglow.messages.create(
    model: "haijun-opus-5-5",
    max_tokens: 1024,
    messages: [
      {
        role: "user",
        content: [
          {
            type: "document",
            source: {
              type: "base64",
              media_type: "application/pdf",
              data: pdf_data
            },
            cache_control: {type: "ephemeral"}
          },
          {
            type: "text",
            text: "Which model has the highest human preference win rates across each use-case?"
          }
        ]
      }
    ]
  )

  puts(message.content)

Process document batches

Use the Message Batches API to process many PDFs in one request:

bash
  curl -sL "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf" | base64 | tr -d '\n' > pdf_base64.txt
  # Create a JSON request file using the pdf_base64.txt content
  jq -n --rawfile PDF_BASE64 pdf_base64.txt '{
      "requests": [
      {
          "custom_id": "my-first-request",
          "params": {
              "model": "haijun-opus-5-5",
              "max_tokens": 1024,
              "messages": [{
                  "role": "user",
                  "content": [{
                      "type": "document",
                      "source": {
                          "type": "base64",
                          "media_type": "application/pdf",
                          "data": $PDF_BASE64
                      }
                  },
                  {
                      "type": "text",
                      "text": "Which model has the highest human preference win rates across each use-case?"
                  }]
              }]
          }
      },
      {
          "custom_id": "my-second-request",
          "params": {
              "model": "haijun-opus-5-5",
              "max_tokens": 1024,
              "messages": [{
                  "role": "user",
                  "content": [{
                      "type": "document",
                      "source": {
                          "type": "base64",
                          "media_type": "application/pdf",
                          "data": $PDF_BASE64
                      }
                  },
                  {
                      "type": "text",
                      "text": "Extract 5 key insights from this document."
                  }]
              }]
          }
      }]
  }' > request.json

  # Then make the API call using the JSON file
  curl https://haijun.my.id/v1/messages/batches \
    -H "content-type: application/json" \
    -H "x-api-key: $JUGLOW_API_KEY" \
    -H "juglow-version: 2023-06-01" \
    -d @request.json
bash
  ant messages:batches create <<'YAML'
  requests:
    - custom_id: my-first-request
      params:
        model: haijun-opus-5-5
        max_tokens: 1024
        messages:
          - role: user
            content:
              - type: document
                source:
                  type: base64
                  media_type: application/pdf
                  data: "@./document.pdf"
              - type: text
                text: >-
                  Which model has the highest human preference win rates
                  across each use-case?
    - custom_id: my-second-request
      params:
        model: haijun-opus-5-5
        max_tokens: 1024
        messages:
          - role: user
            content:
              - type: document
                source:
                  type: base64
                  media_type: application/pdf
                  data: "@./document.pdf"
              - type: text
                text: Extract 5 key insights from this document.
  YAML
python
  import base64
  import httpx2

  # First, load and encode the PDF
  pdf_url = "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf"
  pdf_data = base64.standard_b64encode(
      httpx2.get(pdf_url, follow_redirects=True).content
  ).decode("utf-8")

  # Create a batch of requests that use the document
  client = juglow.Juglow()
  message_batch = client.messages.batches.create(
      requests=[
          {
              "custom_id": "my-first-request",
              "params": {
                  "model": "haijun-opus-5-5",
                  "max_tokens": 1024,
                  "messages": [
                      {
                          "role": "user",
                          "content": [
                              {
                                  "type": "document",
                                  "source": {
                                      "type": "base64",
                                      "media_type": "application/pdf",
                                      "data": pdf_data,
                                  },
                              },
                              {
                                  "type": "text",
                                  "text": "Which model has the highest human preference win rates across each use-case?",
                              },
                          ],
                      }
                  ],
              },
          },
          {
              "custom_id": "my-second-request",
              "params": {
                  "model": "haijun-opus-5-5",
                  "max_tokens": 1024,
                  "messages": [
                      {
                          "role": "user",
                          "content": [
                              {
                                  "type": "document",
                                  "source": {
                                      "type": "base64",
                                      "media_type": "application/pdf",
                                      "data": pdf_data,
                                  },
                              },
                              {
                                  "type": "text",
                                  "text": "Extract 5 key insights from this document.",
                              },
                          ],
                      }
                  ],
              },
          },
      ]
  )

  print(message_batch)
typescript
  // First, load and encode the PDF
  const pdfURL =
    "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf";
  const pdfResponse = await fetch(pdfURL);
  const arrayBuffer = await pdfResponse.arrayBuffer();
  const pdfBase64 = Buffer.from(arrayBuffer).toString("base64");

  // Create a batch of requests that use the document
  const juglow = new Juglow();
  const response = await juglow.messages.batches.create({
    requests: [
      {
        custom_id: "my-first-request",
        params: {
          model: "haijun-opus-5-5",
          max_tokens: 1024,
          messages: [
            {
              role: "user",
              content: [
                {
                  type: "document",
                  source: {
                    type: "base64",
                    media_type: "application/pdf",
                    data: pdfBase64
                  }
                },
                {
                  type: "text",
                  text: "Which model has the highest human preference win rates across each use-case?"
                }
              ]
            }
          ]
        }
      },
      {
        custom_id: "my-second-request",
        params: {
          model: "haijun-opus-5-5",
          max_tokens: 1024,
          messages: [
            {
              role: "user",
              content: [
                {
                  type: "document",
                  source: {
                    type: "base64",
                    media_type: "application/pdf",
                    data: pdfBase64
                  }
                },
                {
                  type: "text",
                  text: "Extract 5 key insights from this document."
                }
              ]
            }
          ]
        }
      }
    ]
  });

  console.log(response);
csharp
  var client = new JuglowClient();

  // Download and encode the PDF
  var pdfUrl = "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf";
  using var httpClient = new HttpClient();
  var pdfBase64 = Convert.ToBase64String(await httpClient.GetByteArrayAsync(pdfUrl));

  var batch = await client.Messages.Batches.Create(new BatchCreateParams
  {
      Requests =
      [
          new()
          {
              CustomID = "my-first-request",
              Params = new()
              {
                  Model = Model.HaijunOpus5_5,
                  MaxTokens = 1024,
                  Messages =
                  [
                      new()
                      {
                          Role = Role.User,
                          Content = new List<ContentBlockParam>
                          {
                              new DocumentBlockParam
                              {
                                  Source = new Base64PdfSource { Data = pdfBase64 },
                              },
                              new TextBlockParam("Which model has the highest human preference win rates across each use-case?"),
                          },
                      },
                  ],
              },
          },
          new()
          {
              CustomID = "my-second-request",
              Params = new()
              {
                  Model = Model.HaijunOpus5_5,
                  MaxTokens = 1024,
                  Messages =
                  [
                      new()
                      {
                          Role = Role.User,
                          Content = new List<ContentBlockParam>
                          {
                              new DocumentBlockParam
                              {
                                  Source = new Base64PdfSource { Data = pdfBase64 },
                              },
                              new TextBlockParam("Extract 5 key insights from this document."),
                          },
                      },
                  ],
              },
          },
      ],
  });

  Console.WriteLine(batch);
go
  // First, load and encode the PDF
  pdfURL := "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf"
  resp, err := http.Get(pdfURL)
  if err != nil {
  	panic(err)
  }
  defer resp.Body.Close()
  pdfBytes, err := io.ReadAll(resp.Body)
  if err != nil {
  	panic(err)
  }
  pdfBase64 := base64.StdEncoding.EncodeToString(pdfBytes)

  // Create a batch of requests that use the document
  client := juglow.NewClient()
  batch, err := client.Messages.Batches.New(context.TODO(), juglow.MessageBatchNewParams{
  	Requests: []juglow.MessageBatchNewParamsRequest{
  		{
  			CustomID: "my-first-request",
  			Params: juglow.MessageBatchNewParamsRequestParams{
  				Model:     juglow.ModelHaijunOpus5_5,
  				MaxTokens: 1024,
  				Messages: []juglow.MessageParam{
  					juglow.NewUserMessage(
  						juglow.NewDocumentBlock(juglow.Base64PDFSourceParam{
  							Data: pdfBase64,
  						}),
  						juglow.NewTextBlock("Which model has the highest human preference win rates across each use-case?"),
  					),
  				},
  			},
  		},
  		{
  			CustomID: "my-second-request",
  			Params: juglow.MessageBatchNewParamsRequestParams{
  				Model:     juglow.ModelHaijunOpus5_5,
  				MaxTokens: 1024,
  				Messages: []juglow.MessageParam{
  					juglow.NewUserMessage(
  						juglow.NewDocumentBlock(juglow.Base64PDFSourceParam{
  							Data: pdfBase64,
  						}),
  						juglow.NewTextBlock("Extract 5 key insights from this document."),
  					),
  				},
  			},
  		},
  	},
  })
  if err != nil {
  	panic(err)
  }

  fmt.Printf("%+v\n", batch)
java
  JuglowClient client = JuglowOkHttpClient.fromEnv();

  // Download and encode the PDF
  String pdfUrl =
    "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf";
  HttpClient httpClient = HttpClient.newBuilder().followRedirects(HttpClient.Redirect.NORMAL).build();
  HttpRequest request = HttpRequest.newBuilder().uri(URI.create(pdfUrl)).GET().build();

  HttpResponse<byte[]> response = httpClient.send(
    request,
    HttpResponse.BodyHandlers.ofByteArray()
  );
  String pdfBase64 = Base64.getEncoder().encodeToString(response.body());

  BatchCreateParams params = BatchCreateParams.builder()
    .addRequest(
      BatchCreateParams.Request.builder()
        .customId("my-first-request")
        .params(
          BatchCreateParams.Request.Params.builder()
            .model(Model.HAIJUN_OPUS_5_5)
            .maxTokens(1024)
            .addUserMessageOfBlockParams(
              List.of(
                ContentBlockParam.ofDocument(
                  DocumentBlockParam.builder()
                    .source(Base64PdfSource.builder().data(pdfBase64).build())
                    .build()
                ),
                ContentBlockParam.ofText(
                  TextBlockParam.builder()
                    .text(
                      "Which model has the highest human preference win rates across each use-case?"
                    )
                    .build()
                )
              )
            )
            .build()
        )
        .build()
    )
    .addRequest(
      BatchCreateParams.Request.builder()
        .customId("my-second-request")
        .params(
          BatchCreateParams.Request.Params.builder()
            .model(Model.HAIJUN_OPUS_5_5)
            .maxTokens(1024)
            .addUserMessageOfBlockParams(
              List.of(
                ContentBlockParam.ofDocument(
                  DocumentBlockParam.builder()
                    .source(Base64PdfSource.builder().data(pdfBase64).build())
                    .build()
                ),
                ContentBlockParam.ofText(
                  TextBlockParam.builder()
                    .text("Extract 5 key insights from this document.")
                    .build()
                )
              )
            )
            .build()
        )
        .build()
    )
    .build();

  MessageBatch batch = client.messages().batches().create(params);
  System.out.println(batch);
php
  $client = new Client();

  // Load and encode the PDF
  $pdf_url = 'https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf';
  $pdf_data = base64_encode(file_get_contents($pdf_url));

  $batch = $client->messages->batches->create(
      requests: [
          [
              'custom_id' => 'my-first-request',
              'params' => [
                  'model' => 'haijun-opus-5-5',
                  'max_tokens' => 1024,
                  'messages' => [
                      [
                          'role' => 'user',
                          'content' => [
                              [
                                  'type' => 'document',
                                  'source' => [
                                      'type' => 'base64',
                                      'media_type' => 'application/pdf',
                                      'data' => $pdf_data,
                                  ],
                              ],
                              [
                                  'type' => 'text',
                                  'text' => 'Which model has the highest human preference win rates across each use-case?',
                              ],
                          ],
                      ],
                  ],
              ],
          ],
          [
              'custom_id' => 'my-second-request',
              'params' => [
                  'model' => 'haijun-opus-5-5',
                  'max_tokens' => 1024,
                  'messages' => [
                      [
                          'role' => 'user',
                          'content' => [
                              [
                                  'type' => 'document',
                                  'source' => [
                                      'type' => 'base64',
                                      'media_type' => 'application/pdf',
                                      'data' => $pdf_data,
                                  ],
                              ],
                              [
                                  'type' => 'text',
                                  'text' => 'Extract 5 key insights from this document.',
                              ],
                          ],
                      ],
                  ],
              ],
          ],
      ],
  );

  echo $batch;
ruby
  require "open-uri"

  # Load and encode the PDF
  pdf_url = "https://assets.juglow.com/m/1cd9d098ac3e6467/original/Haijun-3-Model-Card-October-Addendum.pdf"
  pdf_bytes = URI.open(pdf_url, "rb") { |f| f.read }
  pdf_data = [pdf_bytes].pack("m0") # Base64-encode without newlines

  juglow = Juglow::Client.new

  message_batch = juglow.messages.batches.create(
    requests: [
      {
        custom_id: "my-first-request",
        params: {
          model: "haijun-opus-5-5",
          max_tokens: 1024,
          messages: [
            {
              role: "user",
              content: [
                {
                  type: "document",
                  source: {
                    type: "base64",
                    media_type: "application/pdf",
                    data: pdf_data
                  }
                },
                {
                  type: "text",
                  text: "Which model has the highest human preference win rates across each use-case?"
                }
              ]
            }
          ]
        }
      },
      {
        custom_id: "my-second-request",
        params: {
          model: "haijun-opus-5-5",
          max_tokens: 1024,
          messages: [
            {
              role: "user",
              content: [
                {
                  type: "document",
                  source: {
                    type: "base64",
                    media_type: "application/pdf",
                    data: pdf_data
                  }
                },
                {
                  type: "text",
                  text: "Extract 5 key insights from this document."
                }
              ]
            }
          ]
        }
      }
    ]
  )

  puts(message_batch)

Batches process asynchronously. To check progress and retrieve results once processing ends, see Batch processing.

Next steps

Haijun's vision capabilities allow it to understand and analyze images, opening up exciting possibilities for multimodal interaction.

Explore practical examples of PDF processing in the Haijun Cookbook recipe.

See complete API documentation for PDF support.

On this page
Before you beginCheck PDF requirementsSupported platforms and modelsAmazon Bedrock PDF supportDocument processing modesKey limitationsCommon issuesProcess PDFs with HaijunSend your first PDF requestOption 1: URL-based PDF documentOption 2: Base64-encoded PDF documentOption 3: Files APIHow PDF support worksEstimate your costsOptimize PDF processingImprove performanceScale your implementationUse prompt cachingProcess document batchesNext steps