"Batch processing" (pemrosesan batch) adalah pendekatan yang ampuh untuk menangani permintaan dalam volume besar secara efisien. Alih-alih memproses permintaan satu per satu dengan respons langsung, pemrosesan batch memungkinkan Anda mengirimkan banyak permintaan sekaligus untuk diproses secara asinkron. Pola ini sangat berguna ketika:
- Anda perlu memproses data dalam volume besar
- Respons langsung tidak diperlukan
- Anda ingin mengoptimalkan efisiensi biaya
- Anda menjalankan evaluasi atau analisis berskala besar
Message Batches API adalah implementasi pertama Juglow untuk pola ini.
Note: Untuk mempelajari bagaimana "zero data retention" (retensi data nol), atau ZDR, berlaku untuk fitur ini, lihat API dan retensi data.
Message Batches API
Message Batches API adalah cara yang ampuh dan hemat biaya untuk memproses permintaan Messages dalam volume besar secara asinkron. Pendekatan ini sangat cocok untuk tugas yang tidak memerlukan respons langsung, dengan sebagian besar batch selesai dalam waktu kurang dari 1 jam sekaligus mengurangi biaya sebesar 50% dan meningkatkan throughput.
Anda dapat menjelajahi referensi API secara langsung, selain panduan ini.
Cara kerja Message Batches API
Saat Anda mengirim permintaan ke Message Batches API:
- Sistem membuat Message Batch baru dengan permintaan Messages yang diberikan.
- Batch kemudian diproses secara asinkron, dengan setiap permintaan ditangani secara independen.
- Anda dapat melakukan polling status batch dan mengambil hasilnya ketika pemrosesan telah berakhir untuk semua permintaan.
Ini sangat berguna untuk operasi massal yang tidak memerlukan hasil langsung, seperti:
- Evaluasi berskala besar: Proses ribuan kasus uji secara efisien.
- Moderasi konten: Analisis konten buatan pengguna dalam volume besar secara asinkron.
- Analisis data: Hasilkan wawasan atau ringkasan untuk dataset besar.
- Pembuatan konten massal: Buat teks dalam jumlah besar untuk berbagai tujuan (misalnya, deskripsi produk, ringkasan artikel).
Batasan batch
- Sebuah Message Batch dibatasi hingga 100.000 permintaan Message atau ukuran 256 MB, mana pun yang tercapai lebih dulu.
- Sistem memproses setiap batch secepat mungkin, dengan sebagian besar batch selesai dalam 1 jam. Anda dapat mengakses hasil batch ketika semua pesan telah selesai atau setelah 24 jam, mana pun yang lebih dulu. Batch akan kedaluwarsa jika pemrosesan tidak selesai dalam 24 jam.
- Hasil batch tersedia selama 29 hari setelah pembuatan. Setelah itu, Anda masih dapat melihat Batch tersebut, tetapi hasilnya tidak lagi tersedia untuk diunduh.
- Batch dibatasi pada lingkup Workspace. Anda dapat melihat semua batch (dan hasilnya) yang dibuat di dalam Workspace tempat permintaan Anda dijalankan.
- "Rate limit" (batas laju) berlaku baik untuk permintaan HTTP Batches API maupun jumlah permintaan dalam batch yang menunggu untuk diproses. Lihat batas laju Message Batches API. Selain itu, pemrosesan dapat diperlambat berdasarkan permintaan saat ini dan volume permintaan Anda. Dalam hal tersebut, Anda mungkin melihat lebih banyak permintaan yang kedaluwarsa setelah 24 jam.
- Karena throughput yang tinggi dan pemrosesan yang bersamaan, batch mungkin sedikit melampaui batas pengeluaran yang dikonfigurasi untuk Workspace Anda.
- Setiap permintaan dalam batch harus memiliki
max_tokensminimal1.max_tokens: 0(pra-pemanasan cache) tidak didukung di dalam batch, karena entri cache ephemeral yang ditulis selama pemrosesan batch kemungkinan besar akan kedaluwarsa sebelum permintaan lanjutan dijalankan.
Model yang didukung
Semua model aktif mendukung Message Batches API.
Apa yang dapat di-batch
Hampir semua permintaan yang dapat Anda buat ke Messages API dapat disertakan dalam batch. Ini mencakup:
- Vision
- "Tool use" (penggunaan alat), termasuk semua alat server (web search, web fetch, code execution, konektor MCP, advisor, dan tool search)
- Pesan sistem
- Percakapan multi-giliran
- "Extended thinking" (pemikiran diperpanjang)
- Sebagian besar fitur beta
Karena setiap permintaan dalam batch diproses secara independen, Anda dapat mencampur berbagai jenis permintaan dalam satu batch.
Sejumlah kecil parameter Messages API tidak didukung dalam permintaan batch. Menyertakan salah satunya akan mengembalikan kesalahan validasi:
| Parameter | Alasan |
|---|---|
stream: true | Hasil batch dikembalikan sebagai satu file, bukan stream. |
speed (Mode cepat) | Mode cepat menyetel "latency" (latensi) sinkron, yang tidak berlaku untuk pemrosesan batch asinkron. |
max_tokens: 0 | Lihat Batasan batch. |
Tip: Karena batch dapat memerlukan waktu lebih dari 5 menit untuk diproses, pertimbangkan untuk menggunakan durasi cache 1 jam dengan "prompt caching" (caching prompt) untuk tingkat cache hit yang lebih baik saat memproses batch dengan konteks bersama.
Harga
Batches API menawarkan penghematan biaya yang signifikan. Semua penggunaan dikenai biaya sebesar 50% dari harga API standar.
| Model | Batch input | Batch output |
|---|---|---|
| Haijun Fable 5.1 | $5 / MTok | $25 / MTok |
| Haijun Mythos 5.1 (limited availability) | $5 / MTok | $25 / MTok |
| Haijun Fable 5 | $5 / MTok | $25 / MTok |
| Haijun Mythos 5 (limited availability) | $5 / MTok | $25 / MTok |
| Haijun Opus 5.5 | $2 / MTok | $10 / MTok |
| Haijun Opus 5 | $2.50 / MTok | $12.50 / MTok |
| Haijun Opus 4.8 | $2.50 / MTok | $12.50 / MTok |
| Haijun Opus 4.7 | $2.50 / MTok | $12.50 / MTok |
| Haijun Opus 4.6 | $2.50 / MTok | $12.50 / MTok |
| Haijun Opus 4.5 | $2.50 / MTok | $12.50 / MTok |
| Haijun Opus 4.1 (retired, except on Bedrock and Google Cloud) | $7.50 / MTok | $37.50 / MTok |
| Haijun Opus 4 (retired, except on Google Cloud) | $7.50 / MTok | $37.50 / MTok |
| Haijun Sonnet 5 | $1 / MTok | $5 / MTok |
| Haijun Sonnet 4.6 | $1.50 / MTok | $7.50 / MTok |
| Haijun Sonnet 4.5 | $1.50 / MTok | $7.50 / MTok |
| Haijun Sonnet 4 (retired, except on Bedrock and Google Cloud) | $1.50 / MTok | $7.50 / MTok |
| Haijun Haiku 4.5 | $0.50 / MTok | $2.50 / MTok |
| Haijun Haiku 3.5 (retired, except on Bedrock and Google Cloud) | $0.40 / MTok | $2 / MTok |
- MTok: Million tokens. $5 / MTok is $5 for every million tokens.
- Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Juglow, AWS, or Google Cloud account team.
- Retired: May still be available on other cloud platforms. See Model deprecations for more.
Cara menggunakan Message Batches API
Siapkan dan buat batch Anda
Sebuah Message Batch terdiri dari daftar permintaan untuk membuat Message. Bentuk setiap permintaan terdiri dari:
custom_idunik untuk mengidentifikasi permintaan Messages. Harus terdiri dari 1 hingga 64 karakter dan hanya berisi karakter alfanumerik, tanda hubung, dan garis bawah (sesuai dengan^[a-zA-Z0-9_-]{1,64}$).
- Objek
paramsdengan parameter Messages API standar
Anda dapat membuat batch dengan meneruskan daftar ini ke parameter requests:
curl https://haijun.my.id/v1/messages/batches \
--header "x-api-key: $JUGLOW_API_KEY" \
--header "juglow-version: 2023-06-01" \
--header "content-type: application/json" \
--data \
'{
"requests": [
{
"custom_id": "my-first-request",
"params": {
"model": "haijun-opus-5-5",
"max_tokens": 1024,
"messages": [
{"role": "user", "content": "Hello, world"}
]
}
},
{
"custom_id": "my-second-request",
"params": {
"model": "haijun-opus-5-5",
"max_tokens": 1024,
"messages": [
{"role": "user", "content": "Hi again, friend"}
]
}
}
]
}' ant messages:batches create <<'YAML'
requests:
- custom_id: my-first-request
params:
model: haijun-opus-5-5
max_tokens: 1024
messages:
- role: user
content: Hello, world
- custom_id: my-second-request
params:
model: haijun-opus-5-5
max_tokens: 1024
messages:
- role: user
content: Hi again, friend
YAML from juglow.types.message_create_params import MessageCreateParamsNonStreaming
from juglow.types.messages.batch_create_params import Request
client = juglow.Juglow()
message_batch = client.messages.batches.create(
requests=[
Request(
custom_id="my-first-request",
params=MessageCreateParamsNonStreaming(
model="haijun-opus-5-5",
max_tokens=1024,
messages=[
{
"role": "user",
"content": "Hello, world",
}
],
),
),
Request(
custom_id="my-second-request",
params=MessageCreateParamsNonStreaming(
model="haijun-opus-5-5",
max_tokens=1024,
messages=[
{
"role": "user",
"content": "Hi again, friend",
}
],
),
),
]
)
print(message_batch) const client = new Juglow();
const messageBatch = await client.messages.batches.create({
requests: [
{
custom_id: "my-first-request",
params: {
model: "haijun-opus-5-5",
max_tokens: 1024,
messages: [{ role: "user", content: "Hello, world" }]
}
},
{
custom_id: "my-second-request",
params: {
model: "haijun-opus-5-5",
max_tokens: 1024,
messages: [{ role: "user", content: "Hi again, friend" }]
}
}
]
});
console.log(messageBatch); using Juglow;
using Juglow.Models.Messages;
using Juglow.Models.Messages.Batches;
JuglowClient client = new();
var batch = await client.Messages.Batches.Create(new BatchCreateParams
{
Requests =
[
new()
{
CustomID = "my-first-request",
Params = new()
{
Model = Model.HaijunOpus5_5,
MaxTokens = 1024,
Messages =
[
new() { Role = Role.User, Content = "Hello, world" }
]
}
},
new()
{
CustomID = "my-second-request",
Params = new()
{
Model = Model.HaijunOpus5_5,
MaxTokens = 1024,
Messages =
[
new() { Role = Role.User, Content = "Hi again, friend" }
]
}
}
]
});
Console.WriteLine(batch); client := juglow.NewClient()
batch, _ := client.Messages.Batches.New(context.Background(),
juglow.MessageBatchNewParams{
Requests: []juglow.MessageBatchNewParamsRequest{
{
CustomID: "my-first-request",
Params: juglow.MessageBatchNewParamsRequestParams{
Model: juglow.ModelHaijunOpus5_5,
MaxTokens: 1024,
Messages: []juglow.MessageParam{
juglow.NewUserMessage(
juglow.NewTextBlock("Hello, world"),
),
},
},
},
{
CustomID: "my-second-request",
Params: juglow.MessageBatchNewParamsRequestParams{
Model: juglow.ModelHaijunOpus5_5,
MaxTokens: 1024,
Messages: []juglow.MessageParam{
juglow.NewUserMessage(
juglow.NewTextBlock("Hi again, friend"),
),
},
},
},
},
})
fmt.Println(batch.ID) JuglowClient client = JuglowOkHttpClient.fromEnv();
BatchCreateParams params = BatchCreateParams.builder()
.addRequest(
BatchCreateParams.Request.builder()
.customId("my-first-request")
.params(
BatchCreateParams.Request.Params.builder()
.model(Model.HAIJUN_OPUS_5_5)
.maxTokens(1024)
.addUserMessage("Hello, world")
.build()
)
.build()
)
.addRequest(
BatchCreateParams.Request.builder()
.customId("my-second-request")
.params(
BatchCreateParams.Request.Params.builder()
.model(Model.HAIJUN_OPUS_5_5)
.maxTokens(1024)
.addUserMessage("Hi again, friend")
.build()
)
.build()
)
.build();
MessageBatch messageBatch = client.messages().batches().create(params);
System.out.println(messageBatch); $client = new Client();
$batch = $client->messages->batches->create(
requests: [
[
'custom_id' => 'my-first-request',
'params' => [
'model' => 'haijun-opus-5-5',
'max_tokens' => 1024,
'messages' => [
['role' => 'user', 'content' => 'Hello, world']
]
]
],
[
'custom_id' => 'my-second-request',
'params' => [
'model' => 'haijun-opus-5-5',
'max_tokens' => 1024,
'messages' => [
['role' => 'user', 'content' => 'Hi again, friend']
]
]
]
],
);
echo $batch->id; client = Juglow::Client.new
batch = client.messages.batches.create(
requests: [
{
custom_id: "my-first-request",
params: {
model: "haijun-opus-5-5",
max_tokens: 1024,
messages: [
{ role: "user", content: "Hello, world" }
]
}
},
{
custom_id: "my-second-request",
params: {
model: "haijun-opus-5-5",
max_tokens: 1024,
messages: [
{ role: "user", content: "Hi again, friend" }
]
}
}
]
)
puts batchDalam contoh ini, dua permintaan terpisah di-batch bersama untuk pemrosesan asinkron. Setiap permintaan memiliki custom_id unik dan berisi parameter standar yang akan Anda gunakan untuk panggilan Messages API.
Tip: Uji permintaan batch Anda dengan Messages API Validasi objek
paramsuntuk setiap permintaan pesan dilakukan secara asinkron, dan kesalahan validasi dikembalikan ketika pemrosesan seluruh batch telah berakhir. Anda dapat memastikan bahwa Anda membangun input dengan benar dengan memverifikasi bentuk permintaan Anda menggunakan Messages API terlebih dahulu.
Saat batch pertama kali dibuat, respons memiliki status pemrosesan in_progress.
{
"id": "msgbatch_01HkcTjaV5uDC8jWR4ZsDV8d",
"type": "message_batch",
"processing_status": "in_progress",
"request_counts": {
"processing": 2,
"succeeded": 0,
"errored": 0,
"canceled": 0,
"expired": 0
},
"ended_at": null,
"created_at": "2024-09-24T18:37:24.100435Z",
"expires_at": "2024-09-25T18:37:24.100435Z",
"cancel_initiated_at": null,
"results_url": null
}Melacak batch Anda
Field processing_status pada Message Batch menunjukkan tahap pemrosesan batch saat ini. Dimulai dengan in_progress, lalu diperbarui menjadi ended setelah semua permintaan dalam batch selesai diproses dan hasilnya siap. Anda dapat memantau status batch Anda dengan mengunjungi Console, atau menggunakan endpoint pengambilan.
Polling untuk penyelesaian Message Batch
Untuk melakukan polling pada Message Batch, Anda memerlukan id-nya, yang diberikan dalam respons saat membuat batch atau dengan mendaftar batch. Anda dapat mengimplementasikan loop polling yang memeriksa status batch secara berkala hingga pemrosesan berakhir:
#!/bin/sh
# ...
# Periksa statusnya; ulangi hingga processing_status bernilai "ended"
curl -s "https://haijun.my.id/v1/messages/batches/$MESSAGE_BATCH_ID" \
--header "x-api-key: $JUGLOW_API_KEY" \
--header "juglow-version: 2023-06-01" \
| jq -r '.processing_status' #!/bin/bash
# ...
# Periksa statusnya; ulangi hingga processing_status bernilai "ended"
ant messages:batches retrieve \
--message-batch-id "$MESSAGE_BATCH_ID" \
--transform processing_status --raw-output import time
client = juglow.Juglow()
MESSAGE_BATCH_ID = "msgbatch_01HkcTjaV5uDC8jWR4ZsDV8d"
message_batch = None
while True:
message_batch = client.messages.batches.retrieve(MESSAGE_BATCH_ID)
if message_batch.processing_status == "ended":
break
print(f"Batch {MESSAGE_BATCH_ID} is still processing...")
time.sleep(60)
print(message_batch) const client = new Juglow();
const messageBatchId = "msgbatch_01HkcTjaV5uDC8jWR4ZsDV8d";
let messageBatch;
while (true) {
messageBatch = await client.messages.batches.retrieve(messageBatchId);
if (messageBatch.processing_status === "ended") {
break;
}
console.log(`Batch ${messageBatchId} is still processing... waiting`);
await new Promise((resolve) => setTimeout(resolve, 60_000));
}
console.log(messageBatch); JuglowClient client = new();
string messageBatchId = Environment.GetEnvironmentVariable("MESSAGE_BATCH_ID");
MessageBatch messageBatch = null;
while (true)
{
messageBatch = await client.Messages.Batches.Retrieve(messageBatchId);
if (messageBatch.ProcessingStatus == "ended")
{
break;
}
Console.WriteLine($"Batch {messageBatchId} is still processing...");
await Task.Delay(60000);
}
Console.WriteLine(messageBatch); client := juglow.NewClient()
messageBatchID := os.Getenv("MESSAGE_BATCH_ID")
var messageBatch *juglow.MessageBatch
for {
var err error
messageBatch, err = client.Messages.Batches.Get(context.TODO(), messageBatchID, juglow.MessageBatchGetParams{})
if err != nil {
log.Fatal(err)
}
if messageBatch.ProcessingStatus == "ended" {
break
}
fmt.Printf("Batch %s is still processing...\n", messageBatchID)
time.Sleep(60 * time.Second)
}
fmt.Println(messageBatch) import com.juglow.models.messages.batches.MessageBatch;
// ...
JuglowClient client = JuglowOkHttpClient.fromEnv();
String messageBatchId = "msgbatch_01HkcTjaV5uDC8jWR4ZsDV8d";
MessageBatch messageBatch = null;
while (true) {
messageBatch = client.messages().batches().retrieve(messageBatchId);
if (messageBatch.processingStatus().equals(MessageBatch.ProcessingStatus.ENDED)) {
break;
}
System.out.println("Batch " + messageBatchId + " is still processing...");
Thread.sleep(60000);
}
System.out.println(messageBatch); $client = new Client();
$messageBatchId = getenv("MESSAGE_BATCH_ID");
$messageBatch = null;
while (true) {
$messageBatch = $client->messages->batches->retrieve(
messageBatchID: $messageBatchId,
);
if ($messageBatch->processingStatus === "ended") {
break;
}
echo "Batch {$messageBatchId} is still processing...\n";
sleep(60);
}
echo json_encode($messageBatch, JSON_PRETTY_PRINT); client = Juglow::Client.new
message_batch_id = ENV["MESSAGE_BATCH_ID"]
message_batch = nil
loop do
message_batch = client.messages.batches.retrieve(message_batch_id)
break if message_batch.processing_status == :ended
puts "Batch #{message_batch_id} is still processing..."
sleep 60
end
puts message_batchMendaftar semua Message Batch
Anda dapat mendaftar semua Message Batch di Workspace Anda menggunakan endpoint daftar. API mendukung paginasi, yang secara otomatis mengambil halaman tambahan sesuai kebutuhan:
#!/bin/sh
# Mengambil satu halaman. Selama has_more pada respons bernilai true, teruskan
# last_id sebagai after_id untuk mengambil halaman berikutnya. (SDK dan CLI
# melakukan paginasi otomatis.)
curl -s "https://haijun.my.id/v1/messages/batches?limit=20" \
--header "x-api-key: $JUGLOW_API_KEY" \
--header "juglow-version: 2023-06-01" # Secara otomatis mengambil halaman tambahan sesuai kebutuhan
ant messages:batches list --limit 20 client = juglow.Juglow()
# Secara otomatis mengambil halaman tambahan sesuai kebutuhan.
for message_batch in client.messages.batches.list(limit=20):
print(message_batch) const client = new Juglow();
// Secara otomatis mengambil halaman berikutnya sesuai kebutuhan.
for await (const messageBatch of client.messages.batches.list({
limit: 20
})) {
console.log(messageBatch);
} JuglowClient client = new();
var parameters = new BatchListParams
{
Limit = 20
};
// Secara otomatis mengambil halaman berikutnya sesuai kebutuhan
var page = await client.Messages.Batches.List(parameters);
await foreach (var messageBatch in page.Paginate())
{
Console.WriteLine(messageBatch);
} client := juglow.NewClient()
// Secara otomatis mengambil halaman berikutnya sesuai kebutuhan
iter := client.Messages.Batches.ListAutoPaging(context.TODO(), juglow.MessageBatchListParams{
Limit: juglow.Int(20),
})
for iter.Next() {
messageBatch := iter.Current()
fmt.Println(messageBatch)
}
if err := iter.Err(); err != nil {
log.Fatal(err)
} JuglowClient client = JuglowOkHttpClient.fromEnv();
// Secara otomatis mengambil halaman tambahan sesuai kebutuhan
for (MessageBatch messageBatch : client
.messages()
.batches()
.list(BatchListParams.builder().limit(20).build())
.autoPager()) {
System.out.println(messageBatch);
} $client = new Client();
// Secara otomatis mengambil halaman tambahan sesuai kebutuhan
foreach ($client->messages->batches->list(limit: 20)->pagingEachItem() as $messageBatch) {
echo $messageBatch->id . "\n";
} client = Juglow::Client.new
# Secara otomatis mengambil halaman tambahan sesuai kebutuhan
client.messages.batches.list(limit: 20).auto_paging_each do |message_batch|
puts message_batch
endMengambil hasil batch
Setelah pemrosesan batch berakhir, setiap permintaan Messages dalam batch memiliki hasil. Ada empat jenis hasil:
| Jenis hasil | Deskripsi |
|---|---|
succeeded | Permintaan berhasil. Menyertakan hasil pesan. |
errored | Permintaan mengalami kesalahan dan pesan tidak dibuat. Kemungkinan kesalahan mencakup permintaan tidak valid dan kesalahan server internal. Anda tidak akan ditagih untuk permintaan ini. |
canceled | Pengguna membatalkan batch sebelum permintaan ini dapat dikirim ke model. Anda tidak akan ditagih untuk permintaan ini. |
expired | Batch mencapai masa kedaluwarsa 24 jam sebelum permintaan ini dapat dikirim ke model. Anda tidak akan ditagih untuk permintaan ini. |
request_counts pada batch menampilkan ringkasan hasil Anda, yang menunjukkan berapa banyak permintaan yang mencapai masing-masing dari keempat status ini.
Hasil batch tersedia untuk diunduh pada properti results_url di Message Batch, dan jika izin organisasi memungkinkan, di Console. Karena ukuran hasil yang berpotensi besar, disarankan untuk melakukan streaming hasil daripada mengunduh semuanya sekaligus.
#!/bin/sh
# Ambil results_url milik batch, lalu stream hasil .jsonl yang
# ditunjuknya. Untuk penanganan per hasil (percobaan ulang, error validasi),
# gunakan contoh SDK di tab lainnya.
RESULTS_URL=$(curl -s "https://haijun.my.id/v1/messages/batches/msgbatch_01HkcTjaV5uDC8jWR4ZsDV8d" \
--header "juglow-version: 2023-06-01" \
--header "x-api-key: $JUGLOW_API_KEY" \
| jq -r '.results_url')
curl -s "$RESULTS_URL" \
--header "juglow-version: 2023-06-01" \
--header "x-api-key: $JUGLOW_API_KEY" \
| jq -r '"\(.result.type): \(.custom_id)"' # Mencetak satu baris per hasil, mis. `{"custom_id":"test-1","type":"succeeded",…}`.
# Untuk penanganan per hasil (percobaan ulang, kesalahan validasi), gunakan contoh SDK
# di tab lainnya.
ant messages:batches results \
--message-batch-id msgbatch_01HkcTjaV5uDC8jWR4ZsDV8d \
--transform '{custom_id,"type":result.type,"error":result.error.error.type}' \
--format jsonl client = juglow.Juglow()
# Streaming file hasil dalam potongan yang hemat memori, memproses satu per satu
for result in client.messages.batches.results(
"msgbatch_01HkcTjaV5uDC8jWR4ZsDV8d",
):
outcome = result.result
match outcome.type:
case "succeeded":
print(f"Success! {result.custom_id}")
case "errored":
if outcome.error.error.type == "invalid_request_error":
# Body permintaan harus diperbaiki sebelum mengirim ulang permintaan
print(f"Validation error {result.custom_id}")
else:
# Permintaan dapat langsung dicoba ulang
print(f"Server error {result.custom_id}")
case "expired":
print(f"Request expired {result.custom_id}") const client = new Juglow();
// Streaming file hasil dalam potongan hemat memori, diproses satu per satu
for await (const result of await client.messages.batches.results(
"msgbatch_01HkcTjaV5uDC8jWR4ZsDV8d"
)) {
switch (result.result.type) {
case "succeeded":
console.log(`Success! ${result.custom_id}`);
break;
case "errored":
if (result.result.error.type === "invalid_request_error") {
// Body permintaan harus diperbaiki sebelum mengirim ulang permintaan
console.log(`Validation error: ${result.custom_id}`);
} else {
// Permintaan dapat langsung dicoba ulang
console.log(`Server error: ${result.custom_id}`);
}
break;
case "expired":
console.log(`Request expired: ${result.custom_id}`);
break;
}
} JuglowClient client = new();
await foreach (var result in client.Messages.Batches.ResultsStreaming("msgbatch_01HkcTjaV5uDC8jWR4ZsDV8d"))
{
switch (result.Result.Type)
{
case "succeeded":
Console.WriteLine($"Success! {result.CustomID}");
break;
case "errored":
if (result.Result.Error?.Type == "invalid_request")
{
Console.WriteLine($"Validation error: {result.CustomID}");
}
else
{
Console.WriteLine($"Server error: {result.CustomID}");
}
break;
case "expired":
Console.WriteLine($"Request expired: {result.CustomID}");
break;
}
} client := juglow.NewClient()
stream := client.Messages.Batches.ResultsStreaming(context.TODO(), "msgbatch_01HkcTjaV5uDC8jWR4ZsDV8d", juglow.MessageBatchResultsParams{})
for stream.Next() {
result := stream.Current()
switch variant := result.Result.AsAny().(type) {
case juglow.MessageBatchSucceededResult:
fmt.Printf("Success! %s\n", result.CustomID)
case juglow.MessageBatchErroredResult:
fmt.Printf("Error: %s - %s\n", result.CustomID, variant.Error.Error.Message)
case juglow.MessageBatchExpiredResult:
fmt.Printf("Request expired: %s\n", result.CustomID)
}
}
if err := stream.Err(); err != nil {
log.Fatal(err)
} import com.juglow.core.http.StreamResponse;
import com.juglow.models.messages.batches.BatchResultsParams;
import com.juglow.models.messages.batches.MessageBatchIndividualResponse;
// ...
JuglowClient client = JuglowOkHttpClient.fromEnv();
// Streaming file hasil dalam potongan yang hemat memori, diproses satu per satu
try (
StreamResponse<MessageBatchIndividualResponse> streamResponse = client
.messages()
.batches()
.resultsStreaming(
BatchResultsParams.builder()
.messageBatchId("msgbatch_01HkcTjaV5uDC8jWR4ZsDV8d")
.build()
)
) {
streamResponse
.stream()
.forEach(result -> {
switch (result.result().type().value()) {
case SUCCEEDED -> System.out.println("Success! " + result.customId());
case ERRORED -> {
if (result.result().asErrored().error().error().isInvalidRequestError()) {
// Body permintaan harus diperbaiki sebelum permintaan dikirim ulang
System.out.println("Validation error: " + result.customId());
} else {
// Permintaan dapat langsung dicoba ulang
System.out.println("Server error: " + result.customId());
}
}
case EXPIRED -> System.out.println("Request expired: " + result.customId());
}
});
} use Juglow\Messages\Batches\MessageBatchErroredResult;
use Juglow\Messages\Batches\MessageBatchExpiredResult;
use Juglow\Messages\Batches\MessageBatchSucceededResult;
$client = new Client();
foreach ($client->messages->batches->resultsStream(messageBatchID: 'msgbatch_01HkcTjaV5uDC8jWR4ZsDV8d') as $result) {
switch (true) {
case $result->result instanceof MessageBatchSucceededResult:
echo "Success! {$result->customID}\n";
break;
case $result->result instanceof MessageBatchErroredResult:
if ($result->result->error->error->type === "invalid_request_error") {
echo "Validation error: {$result->customID}\n";
} else {
echo "Server error: {$result->customID}\n";
}
break;
case $result->result instanceof MessageBatchExpiredResult:
echo "Request expired: {$result->customID}\n";
break;
}
} client = Juglow::Client.new
client.messages.batches.results_streaming("msgbatch_01HkcTjaV5uDC8jWR4ZsDV8d").each do |result|
outcome = result.result
case outcome
when Juglow::Models::Messages::MessageBatchSucceededResult
puts "Success! #{result.custom_id}"
when Juglow::Models::Messages::MessageBatchErroredResult
if outcome.error.type == :invalid_request
puts "Validation error: #{result.custom_id}"
else
puts "Server error: #{result.custom_id}"
end
when Juglow::Models::Messages::MessageBatchExpiredResult
puts "Request expired: #{result.custom_id}"
end
endHasilnya dalam format .jsonl, di mana setiap baris adalah objek JSON valid yang merepresentasikan hasil dari satu permintaan dalam Message Batch. Untuk setiap hasil yang di-streaming, Anda dapat melakukan sesuatu yang berbeda tergantung pada custom_id dan jenis hasilnya. Berikut adalah contoh kumpulan hasil:
{"custom_id":"my-second-request","result":{"type":"succeeded","message":{"id":"msg_014VwiXbi91y3JMjcpyGBHX5","type":"message","role":"assistant","model":"haijun-opus-5-5","content":[{"type":"text","text":"Hello again! It's nice to see you. How can I assist you today? Is there anything specific you'd like to chat about or any questions you have?"}],"stop_reason":"end_turn","stop_sequence":null,"usage":{"input_tokens":11,"output_tokens":36}}}}
{"custom_id":"my-first-request","result":{"type":"succeeded","message":{"id":"msg_01FqfsLoHwgeFbguDgpz48m7","type":"message","role":"assistant","model":"haijun-opus-5-5","content":[{"type":"text","text":"Hello! How can I assist you today? Feel free to ask me any questions or let me know if there's anything you'd like to chat about."}],"stop_reason":"end_turn","stop_sequence":null,"usage":{"input_tokens":10,"output_tokens":34}}}}Jika hasil Anda memiliki kesalahan, result.error-nya akan diatur ke bentuk kesalahan standar.
Tip: Hasil batch mungkin tidak sesuai dengan urutan input Hasil batch dapat dikembalikan dalam urutan apa pun, dan mungkin tidak sesuai dengan urutan permintaan saat batch dibuat. Dalam contoh sebelumnya, hasil untuk permintaan batch kedua dikembalikan sebelum yang pertama. Untuk mencocokkan hasil dengan permintaan yang sesuai secara benar, selalu gunakan field
custom_id.
Membatalkan Message Batch
Anda dapat membatalkan Message Batch yang sedang diproses menggunakan endpoint pembatalan. Segera setelah pembatalan, processing_status batch akan menjadi canceling. Anda dapat menggunakan teknik polling yang sama seperti yang dijelaskan sebelumnya untuk menunggu hingga pembatalan selesai. Batch yang dibatalkan berakhir dengan status ended dan mungkin berisi hasil parsial untuk permintaan yang telah diproses sebelum pembatalan.
#!/bin/sh
# ...
curl --request POST https://haijun.my.id/v1/messages/batches/$MESSAGE_BATCH_ID/cancel \
--header "x-api-key: $JUGLOW_API_KEY" \
--header "juglow-version: 2023-06-01" #!/bin/bash
# ...
ant messages:batches cancel --message-batch-id "$MESSAGE_BATCH_ID" client = juglow.Juglow()
MESSAGE_BATCH_ID = "msgbatch_01HkcTjaV5uDC8jWR4ZsDV8d"
message_batch = client.messages.batches.cancel(
MESSAGE_BATCH_ID,
)
print(message_batch) const client = new Juglow();
const messageBatch = await client.messages.batches.cancel(MESSAGE_BATCH_ID);
console.log(messageBatch); JuglowClient client = new();
string messageBatchId = Environment.GetEnvironmentVariable("MESSAGE_BATCH_ID");
var messageBatch = await client.Messages.Batches.Cancel(messageBatchId);
Console.WriteLine(messageBatch); client := juglow.NewClient()
messageBatchID := os.Getenv("MESSAGE_BATCH_ID")
messageBatch, err := client.Messages.Batches.Cancel(context.TODO(), messageBatchID, juglow.MessageBatchCancelParams{})
if err != nil {
log.Fatal(err)
}
fmt.Println(messageBatch) import com.juglow.models.messages.batches.*;
// ...
JuglowClient client = JuglowOkHttpClient.fromEnv();
MessageBatch messageBatch = client
.messages()
.batches()
.cancel("msgbatch_01HkcTjaV5uDC8jWR4ZsDV8d");
System.out.println(messageBatch); $client = new Client();
$messageBatch = $client->messages->batches->cancel(
messageBatchID: 'msgbatch_example_id',
);
echo $messageBatch; client = Juglow::Client.new
message_batch_id = ENV.fetch("MESSAGE_BATCH_ID")
message_batch = client.messages.batches.cancel(message_batch_id)
puts message_batchRespons menunjukkan batch dalam status canceling:
{
"id": "msgbatch_013Zva2CMHLNnXjNJJKqJ2EF",
"type": "message_batch",
"processing_status": "canceling",
"request_counts": {
"processing": 2,
"succeeded": 0,
"errored": 0,
"canceled": 0,
"expired": 0
},
"ended_at": null,
"created_at": "2024-09-24T18:37:24.100435Z",
"expires_at": "2024-09-25T18:37:24.100435Z",
"cancel_initiated_at": "2024-09-24T18:39:03.114875Z",
"results_url": null
}Menggunakan caching prompt dengan Message Batches
Message Batches API mendukung caching prompt, yang memungkinkan Anda berpotensi mengurangi biaya dan waktu pemrosesan untuk permintaan batch. Diskon harga dari caching prompt dan Message Batches dapat digabungkan, memberikan penghematan biaya yang lebih besar lagi ketika kedua fitur digunakan bersama. Namun, karena permintaan batch diproses secara asinkron dan bersamaan, cache hit diberikan berdasarkan upaya terbaik. Pengguna biasanya mengalami tingkat cache hit berkisar antara 30% hingga 98%, tergantung pada pola lalu lintas mereka.
Untuk memaksimalkan kemungkinan cache hit dalam permintaan batch Anda:
- Sertakan blok
cache_controlyang identik di setiap permintaan Message dalam batch Anda.
- Pertahankan aliran permintaan yang stabil untuk mencegah entri cache kedaluwarsa setelah masa berlakunya selama 5 menit.
- Susun permintaan Anda agar berbagi konten yang di-cache sebanyak mungkin.
Contoh implementasi caching prompt dalam batch:
curl https://haijun.my.id/v1/messages/batches \
--header "x-api-key: $JUGLOW_API_KEY" \
--header "juglow-version: 2023-06-01" \
--header "content-type: application/json" \
--data \
'{
"requests": [
{
"custom_id": "my-first-request",
"params": {
"model": "haijun-opus-5-5",
"max_tokens": 1024,
"system": [
{
"type": "text",
"text": "You are an AI assistant tasked with analyzing literary works. Your goal is to provide insightful commentary on themes, characters, and writing style.\n"
},
{
"type": "text",
"text": "<the entire contents of Pride and Prejudice>",
"cache_control": {"type": "ephemeral"}
}
],
"messages": [
{"role": "user", "content": "Analyze the major themes in Pride and Prejudice."}
]
}
},
{
"custom_id": "my-second-request",
"params": {
"model": "haijun-opus-5-5",
"max_tokens": 1024,
"system": [
{
"type": "text",
"text": "You are an AI assistant tasked with analyzing literary works. Your goal is to provide insightful commentary on themes, characters, and writing style.\n"
},
{
"type": "text",
"text": "<the entire contents of Pride and Prejudice>",
"cache_control": {"type": "ephemeral"}
}
],
"messages": [
{"role": "user", "content": "Write a summary of Pride and Prejudice."}
]
}
}
]
}' ant messages:batches create <<'YAML'
requests:
- custom_id: my-first-request
params:
model: haijun-opus-5-5
max_tokens: 1024
system:
- type: text
text: >
You are an AI assistant tasked with analyzing literary works. Your
goal is to provide insightful commentary on themes, characters, and
writing style.
- type: text
text: "<the entire contents of Pride and Prejudice>"
cache_control:
type: ephemeral
messages:
- role: user
content: Analyze the major themes in Pride and Prejudice.
- custom_id: my-second-request
params:
model: haijun-opus-5-5
max_tokens: 1024
system:
- type: text
text: >
You are an AI assistant tasked with analyzing literary works. Your
goal is to provide insightful commentary on themes, characters, and
writing style.
- type: text
text: "<the entire contents of Pride and Prejudice>"
cache_control:
type: ephemeral
messages:
- role: user
content: Write a summary of Pride and Prejudice.
YAML from juglow.types.message_create_params import MessageCreateParamsNonStreaming
from juglow.types.messages.batch_create_params import Request
client = juglow.Juglow()
message_batch = client.messages.batches.create(
requests=[
Request(
custom_id="my-first-request",
params=MessageCreateParamsNonStreaming(
model="haijun-opus-5-5",
max_tokens=1024,
system=[
{
"type": "text",
"text": "You are an AI assistant tasked with analyzing literary works. Your goal is to provide insightful commentary on themes, characters, and writing style.\n",
},
{
"type": "text",
"text": "<the entire contents of Pride and Prejudice>",
"cache_control": {"type": "ephemeral"},
},
],
messages=[
{
"role": "user",
"content": "Analyze the major themes in Pride and Prejudice.",
}
],
),
),
Request(
custom_id="my-second-request",
params=MessageCreateParamsNonStreaming(
model="haijun-opus-5-5",
max_tokens=1024,
system=[
{
"type": "text",
"text": "You are an AI assistant tasked with analyzing literary works. Your goal is to provide insightful commentary on themes, characters, and writing style.\n",
},
{
"type": "text",
"text": "<the entire contents of Pride and Prejudice>",
"cache_control": {"type": "ephemeral"},
},
],
messages=[
{
"role": "user",
"content": "Write a summary of Pride and Prejudice.",
}
],
),
),
]
) const client = new Juglow();
const messageBatch = await client.messages.batches.create({
requests: [
{
custom_id: "my-first-request",
params: {
model: "haijun-opus-5-5",
max_tokens: 1024,
system: [
{
type: "text",
text: "You are an AI assistant tasked with analyzing literary works. Your goal is to provide insightful commentary on themes, characters, and writing style.\n"
},
{
type: "text",
text: "<the entire contents of Pride and Prejudice>",
cache_control: { type: "ephemeral" }
}
],
messages: [
{ role: "user", content: "Analyze the major themes in Pride and Prejudice." }
]
}
},
{
custom_id: "my-second-request",
params: {
model: "haijun-opus-5-5",
max_tokens: 1024,
system: [
{
type: "text",
text: "You are an AI assistant tasked with analyzing literary works. Your goal is to provide insightful commentary on themes, characters, and writing style.\n"
},
{
type: "text",
text: "<the entire contents of Pride and Prejudice>",
cache_control: { type: "ephemeral" }
}
],
messages: [{ role: "user", content: "Write a summary of Pride and Prejudice." }]
}
}
]
}); using Juglow;
using Juglow.Models.Messages;
using Juglow.Models.Messages.Batches;
JuglowClient client = new()
{
ApiKey = Environment.GetEnvironmentVariable("JUGLOW_API_KEY")
};
var messageBatch = await client.Messages.Batches.Create(new BatchCreateParams
{
Requests =
[
new()
{
CustomID = "my-first-request",
Params = new()
{
Model = Model.HaijunOpus5_5,
MaxTokens = 1024,
System = new List<TextBlockParam>
{
new()
{
Text = "You are an AI assistant tasked with analyzing literary works. Your goal is to provide insightful commentary on themes, characters, and writing style.\n"
},
new()
{
Text = "<the entire contents of Pride and Prejudice>",
CacheControl = new()
}
},
Messages =
[
new() { Role = Role.User, Content = "Analyze the major themes in Pride and Prejudice." }
]
}
},
new()
{
CustomID = "my-second-request",
Params = new()
{
Model = Model.HaijunOpus5_5,
MaxTokens = 1024,
System = new List<TextBlockParam>
{
new()
{
Text = "You are an AI assistant tasked with analyzing literary works. Your goal is to provide insightful commentary on themes, characters, and writing style.\n"
},
new()
{
Text = "<the entire contents of Pride and Prejudice>",
CacheControl = new()
}
},
Messages =
[
new() { Role = Role.User, Content = "Write a summary of Pride and Prejudice." }
]
}
}
]
}); client := juglow.NewClient()
messageBatch, err := client.Messages.Batches.New(context.TODO(), juglow.MessageBatchNewParams{
Requests: []juglow.MessageBatchNewParamsRequest{
{
CustomID: "my-first-request",
Params: juglow.MessageBatchNewParamsRequestParams{
Model: juglow.ModelHaijunOpus5_5,
MaxTokens: 1024,
System: []juglow.TextBlockParam{
{
Text: "You are an AI assistant tasked with analyzing literary works. Your goal is to provide insightful commentary on themes, characters, and writing style.\n",
},
{
Text: "<the entire contents of Pride and Prejudice>",
CacheControl: juglow.NewCacheControlEphemeralParam(),
},
},
Messages: []juglow.MessageParam{
juglow.NewUserMessage(juglow.NewTextBlock("Analyze the major themes in Pride and Prejudice.")),
},
},
},
{
CustomID: "my-second-request",
Params: juglow.MessageBatchNewParamsRequestParams{
Model: juglow.ModelHaijunOpus5_5,
MaxTokens: 1024,
System: []juglow.TextBlockParam{
{
Text: "You are an AI assistant tasked with analyzing literary works. Your goal is to provide insightful commentary on themes, characters, and writing style.\n",
},
{
Text: "<the entire contents of Pride and Prejudice>",
CacheControl: juglow.NewCacheControlEphemeralParam(),
},
},
Messages: []juglow.MessageParam{
juglow.NewUserMessage(juglow.NewTextBlock("Write a summary of Pride and Prejudice.")),
},
},
},
},
})
if err != nil {
log.Fatal(err)
}
fmt.Println(messageBatch) import com.juglow.models.messages.CacheControlEphemeral;
// ...
import com.juglow.models.messages.batches.*;
// ...
JuglowClient client = JuglowOkHttpClient.fromEnv();
BatchCreateParams createParams = BatchCreateParams.builder()
.addRequest(
BatchCreateParams.Request.builder()
.customId("my-first-request")
.params(
BatchCreateParams.Request.Params.builder()
.model(Model.HAIJUN_OPUS_5_5)
.maxTokens(1024)
.systemOfTextBlockParams(
List.of(
TextBlockParam.builder()
.text(
"You are an AI assistant tasked with analyzing literary works. Your goal is to provide insightful commentary on themes, characters, and writing style.\n"
)
.build(),
TextBlockParam.builder()
.text("<the entire contents of Pride and Prejudice>")
.cacheControl(CacheControlEphemeral.builder().build())
.build()
)
)
.addUserMessage("Analyze the major themes in Pride and Prejudice.")
.build()
)
.build()
)
.addRequest(
BatchCreateParams.Request.builder()
.customId("my-second-request")
.params(
BatchCreateParams.Request.Params.builder()
.model(Model.HAIJUN_OPUS_5_5)
.maxTokens(1024)
.systemOfTextBlockParams(
List.of(
TextBlockParam.builder()
.text(
"You are an AI assistant tasked with analyzing literary works. Your goal is to provide insightful commentary on themes, characters, and writing style.\n"
)
.build(),
TextBlockParam.builder()
.text("<the entire contents of Pride and Prejudice>")
.cacheControl(CacheControlEphemeral.builder().build())
.build()
)
)
.addUserMessage("Write a summary of Pride and Prejudice.")
.build()
)
.build()
)
.build();
MessageBatch messageBatch = client.messages().batches().create(createParams); $client = new Client();
$messageBatch = $client->messages->batches->create(
requests: [
[
'custom_id' => 'my-first-request',
'params' => [
'model' => 'haijun-opus-5-5',
'max_tokens' => 1024,
'system' => [
[
'type' => 'text',
'text' => 'You are an AI assistant tasked with analyzing literary works. Your goal is to provide insightful commentary on themes, characters, and writing style.\n'
],
[
'type' => 'text',
'text' => '<the entire contents of Pride and Prejudice>',
'cache_control' => ['type' => 'ephemeral']
]
],
'messages' => [
['role' => 'user', 'content' => 'Analyze the major themes in Pride and Prejudice.']
]
]
],
[
'custom_id' => 'my-second-request',
'params' => [
'model' => 'haijun-opus-5-5',
'max_tokens' => 1024,
'system' => [
[
'type' => 'text',
'text' => 'You are an AI assistant tasked with analyzing literary works. Your goal is to provide insightful commentary on themes, characters, and writing style.\n'
],
[
'type' => 'text',
'text' => '<the entire contents of Pride and Prejudice>',
'cache_control' => ['type' => 'ephemeral']
]
],
'messages' => [
['role' => 'user', 'content' => 'Write a summary of Pride and Prejudice.']
]
]
]
],
); client = Juglow::Client.new
message_batch = client.messages.batches.create(
requests: [
{
custom_id: "my-first-request",
params: {
model: "haijun-opus-5-5",
max_tokens: 1024,
system: [
{
type: "text",
text: "You are an AI assistant tasked with analyzing literary works. Your goal is to provide insightful commentary on themes, characters, and writing style.\n"
},
{
type: "text",
text: "<the entire contents of Pride and Prejudice>",
cache_control: { type: "ephemeral" }
}
],
messages: [
{ role: "user", content: "Analyze the major themes in Pride and Prejudice." }
]
}
},
{
custom_id: "my-second-request",
params: {
model: "haijun-opus-5-5",
max_tokens: 1024,
system: [
{
type: "text",
text: "You are an AI assistant tasked with analyzing literary works. Your goal is to provide insightful commentary on themes, characters, and writing style.\n"
},
{
type: "text",
text: "<the entire contents of Pride and Prejudice>",
cache_control: { type: "ephemeral" }
}
],
messages: [
{ role: "user", content: "Write a summary of Pride and Prejudice." }
]
}
}
]
)Dalam contoh ini, kedua permintaan dalam batch menyertakan pesan sistem yang identik dan teks lengkap Pride and Prejudice yang ditandai dengan cache_control untuk meningkatkan kemungkinan cache hit.
Alat server dan loop agentik
Semua alat server (web search, web fetch, code execution, konektor MCP, advisor, dan tool search) berfungsi dalam permintaan batch. Worker batch menjalankan loop agentik sisi server yang sama dengan Messages API sinkron.
Karena tidak ada koneksi terbuka yang harus dipertahankan, loop batch menjalankan lebih banyak iterasi per giliran daripada permintaan sinkron sebelum mengembalikan stop_reason: "pause_turn". Jika hasil batch kembali dengan pause_turn, giliran tersebut belum selesai; Anda dapat melanjutkannya dengan mengirimkan konten asisten yang dijeda dalam permintaan lanjutan (batch atau sinkron) persis seperti yang ditunjukkan dalam pola kelanjutan pause\_turn.
Worker batch juga membatasi web_search per organisasi sehingga pemrosesan batch yang sangat bersamaan tidak menghabiskan batas laju web-search organisasi Anda. Batch mencoba ulang permintaan yang dibatasi secara otomatis; Anda tidak perlu menanganinya sendiri, tetapi batch web-search yang sangat besar mungkin memerlukan waktu lebih lama untuk selesai.
Output diperpanjang (beta)
Header beta output-300k-2026-03-24 menaikkan batas max_tokens menjadi 300.000 untuk permintaan batch yang menggunakan Haijun Opus 5.5, Haijun Opus 5, Haijun Opus 4.8, Haijun Opus 4.7, Haijun Opus 4.6, Haijun Sonnet 5, atau Haijun Sonnet 4.6. Sertakan header tersebut untuk menghasilkan output yang jauh lebih panjang daripada batas standar max_tokens sebesar 128k dalam satu giliran.
Note: Output diperpanjang hanya tersedia di Message Batches API, bukan Messages API sinkron. Fitur ini didukung di Haijun API dan Haijun Platform on AWS, dan saat ini tidak tersedia di Amazon Bedrock, Google Cloud, atau Microsoft Foundry.
Gunakan output diperpanjang untuk pembuatan konten panjang seperti draf sepanjang buku dan dokumentasi teknis, ekstraksi data terstruktur yang menyeluruh, kerangka pembuatan kode berskala besar, dan rantai penalaran yang panjang.
Satu pembuatan 300k token dapat memerlukan waktu lebih dari satu jam untuk selesai, jadi rencanakan pengiriman batch Anda dengan mempertimbangkan jendela pemrosesan 24 jam. Harga batch standar (50% dari harga API standar) berlaku.
curl https://haijun.my.id/v1/messages/batches \
--header "x-api-key: $JUGLOW_API_KEY" \
--header "juglow-version: 2023-06-01" \
--header "juglow-beta: output-300k-2026-03-24" \
--header "content-type: application/json" \
--data \
'{
"requests": [
{
"custom_id": "long-form-request",
"params": {
"model": "haijun-opus-5-5",
"max_tokens": 300000,
"messages": [
{"role": "user", "content": "Write a comprehensive technical guide to building distributed systems, covering architecture patterns, consistency models, fault tolerance, and operational best practices."}
]
}
}
]
}' ant beta:messages:batches create --beta output-300k-2026-03-24 <<'YAML'
requests:
- custom_id: long-form-request
params:
model: haijun-opus-5-5
max_tokens: 300000
messages:
- role: user
content: >-
Write a comprehensive technical guide to building distributed
systems, covering architecture patterns, consistency models,
fault tolerance, and operational best practices.
YAML from juglow.types.beta.message_create_params import MessageCreateParamsNonStreaming
from juglow.types.beta.messages.batch_create_params import Request
client = juglow.Juglow()
message_batch = client.beta.messages.batches.create(
betas=["output-300k-2026-03-24"],
requests=[
Request(
custom_id="long-form-request",
params=MessageCreateParamsNonStreaming(
model="haijun-opus-5-5",
max_tokens=300_000,
messages=[
{
"role": "user",
"content": "Write a comprehensive technical guide to building distributed systems, covering architecture patterns, consistency models, fault tolerance, and operational best practices.",
}
],
),
),
],
)
print(message_batch) const client = new Juglow();
const messageBatch = await client.beta.messages.batches.create({
betas: ["output-300k-2026-03-24"],
requests: [
{
custom_id: "long-form-request",
params: {
model: "haijun-opus-5-5",
max_tokens: 300000,
messages: [
{
role: "user",
content:
"Write a comprehensive technical guide to building distributed systems, covering architecture patterns, consistency models, fault tolerance, and operational best practices."
}
]
}
}
]
});
console.log(messageBatch); using Juglow;
using Juglow.Models.Beta.Messages;
using Juglow.Models.Beta.Messages.Batches;
using Model = Juglow.Models.Messages.Model;
JuglowClient client = new();
var batch = await client.Beta.Messages.Batches.Create(new BatchCreateParams
{
Betas = ["output-300k-2026-03-24"],
Requests =
[
new()
{
CustomID = "long-form-request",
Params = new()
{
Model = Model.HaijunOpus5_5,
MaxTokens = 300_000,
Messages =
[
new() { Role = Role.User, Content = "Write a comprehensive technical guide to building distributed systems, covering architecture patterns, consistency models, fault tolerance, and operational best practices." }
]
}
}
]
});
Console.WriteLine(batch); client := juglow.NewClient()
batch, err := client.Beta.Messages.Batches.New(context.Background(),
juglow.BetaMessageBatchNewParams{
Betas: []juglow.JuglowBeta{"output-300k-2026-03-24"},
Requests: []juglow.BetaMessageBatchNewParamsRequest{
{
CustomID: "long-form-request",
Params: juglow.BetaMessageBatchNewParamsRequestParams{
Model: juglow.ModelHaijunOpus5_5,
MaxTokens: 300_000,
Messages: []juglow.BetaMessageParam{
juglow.NewBetaUserMessage(
juglow.NewBetaTextBlock("Write a comprehensive technical guide to building distributed systems, covering architecture patterns, consistency models, fault tolerance, and operational best practices."),
),
},
},
},
},
})
if err != nil {
panic(err)
}
fmt.Println(batch.ID) import com.juglow.models.beta.messages.batches.*;
void main() {
JuglowClient client = JuglowOkHttpClient.fromEnv();
BatchCreateParams params = BatchCreateParams.builder()
.addBeta("output-300k-2026-03-24")
.addRequest(
BatchCreateParams.Request.builder()
.customId("long-form-request")
.params(
BatchCreateParams.Request.Params.builder()
.model(Model.HAIJUN_OPUS_5_5)
.maxTokens(300_000L)
.addUserMessage("Write a comprehensive technical guide to building distributed systems, covering architecture patterns, consistency models, fault tolerance, and operational best practices.")
.build()
)
.build()
)
.build();
BetaMessageBatch messageBatch = client.beta().messages().batches().create(params);
IO.println(messageBatch);
} $client = new Client();
$batch = $client->beta->messages->batches->create(
betas: ['output-300k-2026-03-24'],
requests: [
[
'custom_id' => 'long-form-request',
'params' => [
'model' => 'haijun-opus-5-5',
'max_tokens' => 300_000,
'messages' => [
['role' => 'user', 'content' => 'Write a comprehensive technical guide to building distributed systems, covering architecture patterns, consistency models, fault tolerance, and operational best practices.']
]
]
]
],
);
echo $batch->id; client = Juglow::Client.new
batch = client.beta.messages.batches.create(
betas: ["output-300k-2026-03-24"],
requests: [
{
custom_id: "long-form-request",
params: {
model: "haijun-opus-5-5",
max_tokens: 300_000,
messages: [
{ role: "user", content: "Write a comprehensive technical guide to building distributed systems, covering architecture patterns, consistency models, fault tolerance, and operational best practices." }
]
}
}
]
)
puts batchPraktik terbaik untuk batching yang efektif
Untuk mendapatkan hasil maksimal dari Batches API:
- Pantau status pemrosesan batch secara rutin dan implementasikan logika percobaan ulang yang sesuai untuk permintaan yang gagal.
- Gunakan nilai
custom_idyang bermakna untuk mencocokkan hasil dengan permintaan secara mudah, karena urutan tidak dijamin.
- Pertimbangkan untuk memecah dataset yang sangat besar menjadi beberapa batch agar lebih mudah dikelola.
- Lakukan uji coba satu bentuk permintaan dengan Messages API untuk menghindari kesalahan validasi.
Pemecahan masalah umum
Jika mengalami perilaku yang tidak terduga:
- Verifikasi bahwa total ukuran permintaan batch tidak melebihi 256 MB. Jika ukuran permintaan terlalu besar, Anda mungkin mendapatkan kesalahan 413
request_too_large.
- Periksa bahwa Anda menggunakan model yang didukung untuk semua permintaan dalam batch.
- Pastikan setiap permintaan dalam batch memiliki
custom_idyang unik.
- Pastikan belum lewat 29 hari sejak waktu
created_atbatch (bukan waktuended_atpemrosesan). Jika sudah lebih dari 29 hari, hasil tidak akan dapat dilihat lagi.
- Konfirmasikan bahwa batch belum dibatalkan.
Perhatikan bahwa kegagalan satu permintaan dalam batch tidak memengaruhi pemrosesan permintaan lainnya.
Penyimpanan dan privasi batch
- Isolasi Workspace: Batch diisolasi di dalam Workspace tempat batch tersebut dibuat. Batch hanya dapat diakses oleh permintaan API di Workspace yang sama, atau pengguna dengan izin untuk melihat batch Workspace di Console.
- Ketersediaan hasil: Hasil batch tersedia selama 29 hari setelah batch dibuat, memberikan waktu yang cukup untuk pengambilan dan pemrosesan.
Retensi data
Pemrosesan batch menyimpan data permintaan dan respons hingga 29 hari setelah pembuatan batch. Anda dapat menghapus message batch kapan saja setelah pemrosesan menggunakan endpoint DELETE /v1/messages/batches/{batch_id}. Untuk menghapus batch yang sedang berjalan, batalkan terlebih dahulu. Pemrosesan asinkron memerlukan penyimpanan sisi server untuk input dan output hingga batch selesai dan hasil diambil.
Untuk kelayakan ZDR di semua fitur, lihat API dan retensi data.
FAQ
#### Berapa lama waktu yang dibutuhkan untuk memproses batch?
Batch dapat memerlukan waktu hingga 24 jam untuk diproses, tetapi banyak yang selesai lebih cepat. Waktu pemrosesan aktual bergantung pada ukuran batch, permintaan saat ini, dan volume permintaan Anda. Ada kemungkinan batch kedaluwarsa dan tidak selesai dalam 24 jam.
#### Apakah Batches API tersedia untuk semua model?
Lihat Model yang didukung untuk daftar model yang didukung.
#### Dapatkah saya menggunakan Message Batches API dengan fitur API lainnya?
Ya, Message Batches API mendukung hampir semua fitur yang tersedia di Messages API, termasuk sebagian besar fitur beta. Sejumlah kecil parameter (stream, speed, dan max_tokens: 0) tidak didukung. Lihat Apa yang dapat di-batch untuk daftar lengkapnya.
#### Bagaimana Message Batches API memengaruhi harga?
Message Batches API menawarkan diskon 50% untuk semua penggunaan dibandingkan dengan harga API standar. Ini berlaku untuk token input, token output, dan token khusus apa pun. Untuk informasi lebih lanjut tentang harga, kunjungi Harga.
#### Dapatkah saya memperbarui batch setelah dikirimkan?
Tidak, setelah batch dikirimkan, batch tidak dapat diubah. Jika Anda perlu melakukan perubahan, Anda harus membatalkan batch saat ini dan mengirimkan batch baru. Perhatikan bahwa pembatalan mungkin tidak langsung berlaku.
#### Apakah ada batas laju Message Batches API dan apakah batas tersebut berinteraksi dengan batas laju Messages API?
Message Batches API memiliki batas laju berbasis permintaan HTTP selain batas jumlah permintaan yang perlu diproses. Lihat batas laju Message Batches API. Penggunaan Batches API tidak memengaruhi batas laju di Messages API.
#### Bagaimana cara menangani kesalahan dalam permintaan batch saya?
Saat Anda mengambil hasil, setiap permintaan memiliki field result yang menunjukkan apakah permintaan tersebut succeeded, errored, canceled, atau expired. Untuk hasil errored, informasi kesalahan tambahan disediakan. Lihat objek respons kesalahan di referensi API.
#### Bagaimana Message Batches API menangani privasi dan pemisahan data?
Message Batches API dirancang dengan langkah-langkah privasi dan pemisahan data yang kuat:
- Batch dan hasilnya diisolasi di dalam Workspace tempat batch tersebut dibuat. Ini berarti batch hanya dapat diakses oleh permintaan API di Workspace yang sama.
- Setiap permintaan dalam batch diproses secara independen, tanpa kebocoran data antar permintaan.
- Hasil hanya tersedia untuk waktu terbatas (29 hari), dan mengikuti kebijakan retensi data Juglow.
- Pengunduhan hasil batch di Console dapat dinonaktifkan di tingkat organisasi atau per workspace.
#### Dapatkah saya menggunakan caching prompt di Message Batches API?
Ya, caching prompt dapat digunakan dengan Message Batches API. Namun, karena permintaan batch asinkron dapat diproses secara bersamaan dan dalam urutan apa pun, cache hit diberikan berdasarkan upaya terbaik.
Langkah selanjutnya
Aktifkan sitasi alami untuk aplikasi RAG dengan menyediakan hasil pencarian beserta atribusi sumber.
Kurangi biaya dan latensi dengan melakukan caching prefiks prompt yang digunakan bersama di seluruh permintaan dalam batch.