Upstage Launches Solar Pro 4 Model
Upstage launched Solar Pro 4, a closed commercial large language model, on August 14, with a 512K-token context window and up to 128K output tokens. The company positions the release for agentic, document-heavy enterprise work, while Chosun reports it became the first Korean model listed on OpenRouter and Hermes Agent.
Upstage launched Solar Pro 4, its flagship closed commercial large language model, on August 14. The Seoul-based company states that the model supports a 512K-token context window, up to 128K output tokens, and English, Korean, and Japanese input and output.
Solar Pro 4 is available through Upstage Console and OpenRouter. Upstage announced a 90% launch discount on those services through September 10. According to Chosun, the release made Solar Pro 4 the first Korean model listed on both OpenRouter and Nous Research's Hermes Agent platform.
Agent and long-context claims
Upstage describes Solar Pro 4 as a model for multi-step work involving documents and tools, including reading files, invoking tools, and producing a final deliverable. The company states that users can select different reasoning-effort settings, from deeper analysis to lower-latency responses.
The company reported the following evaluation results, citing Artificial Analysis data as of August 2026:
- •57 on Terminal-Bench v2.1, a benchmark for multi-step terminal tasks
- •23 on tau3-Banking, an evaluation of multi-turn tool use
- •71 on AA-LCR, a long-context reasoning benchmark
Chosun reports that these scores represent gains over Solar Pro 3 of 4.8 times on TerminalBench, 2.6 times on Tau3-Banking, and more than 2.3 times on AA-LCR. The publication also reports an overall Artificial Analysis score of 42, three times the prior model's result. The source material does not include independent reproduction of these results.
According to Upstage's launch post, the model was trained using its OfficeVerse synthetic-data pipeline, which the company has used since Solar Open 2. The available source material does not provide sufficient technical detail to assess the training mix, data provenance, or evaluation methodology behind the reported improvements.
Relevance for regulated document workflows
Insurance Innovation Reporter describes Solar Pro 4 as positioned for regulated-industry uses, including insurance workflows. Upstage told the publication that the model is designed for reliable instruction following, policy constraints, valid output schemas, and document-grounded responses. The company also claims the model can decline to verify an answer when the evidence is insufficient, cite supported answers, flag clauses absent from a document, and identify conflicting figures.
Those capabilities matter in document-intensive workflows such as claims processing, underwriting, policy review, and document intake, where unsupported extraction or summarization can create compliance and review burdens. More broadly, teams evaluating agentic systems commonly need to measure task-completion reliability, tool-call error rates, and evidence traceability alongside token pricing. Benchmark scores can be useful screening signals, but production validation generally requires representative documents, controlled tool permissions, schema checks, and human review paths.
Upstage CEO Kim Sung-hun said in comments reported by Chosun, "Being listed on global developer platforms as the first Korean model proves that our AI technology is validated on the world stage."
Key Points
- 1Upstage released Solar Pro 4 with 512K context and 128K output capacity, targeting multi-document and multi-step agent workloads.
- 2Vendor-cited benchmark gains cover terminal tasks, tool use, and long-context reasoning, but independent reproduction and methodology details remain limited.
- 3For regulated AI deployments, industry practice emphasizes evidence traceability, schema validation, and workflow-level error testing beyond benchmark performance.
Scoring Rationale
Solar Pro 4 is a notable commercial model release with unusually large context capacity and reported agent-evaluation improvements. Its relevance is strongest for teams building document-heavy enterprise agents, although the available performance evidence is primarily company-reported and the model is closed.
Sources
Primary source and supporting public references used for this report.
Practice with real Ride-Hailing data
90 SQL & Python problems · 15 industry datasets
250 free problems · No credit card
See all Ride-Hailing problems
