Claude 3 Launch: First Impressions
Anthropic has released the Claude 3 family—Haiku, Sonnet, and Opus—with strong vision capabilities, a 200k-token context window, and an intriguing new take on needle-in-a-haystack testing.
The Claude 3 Lineup
Claude 3 was announced overnight. Its lineup, from smallest to largest, is Haiku, Sonnet, and Opus. They are terms associated with poetry or music, although the intended theme is not immediately clear.
Based on the announced benchmark results, the rough positioning appears to be as follows:
- Opus: roughly GPT-4-class, or Gemini 1.0 Ultra-class
- Sonnet: roughly GPT-3.75-class
- Haiku: roughly GPT-3.5-class, or Gemini 1.0 Pro-class
Benchmark Context and Pricing
The GPT-4 figures in the published benchmarks are older results. GPT-4 Turbo has continued to improve beyond those numbers; for example, its GSM8K result is 95.3%, rather than 92.0%.
A rough comparison of pricing per 1 million tokens, for input/output, is below:
- GPT-4 Turbo: $10/$30
- GPT-3.5 Turbo: $0.5/$1.5
- Claude 3 Opus: $15/$75
- Claude 3 Sonnet: $3/$15
- Claude 3 Haiku: $0.25/$1.25
Early Takeaways
Opus is quite expensive. Unless it excels in a particularly distinctive area, it may be difficult to justify using it regularly.
Sonnet offers mid-range performance at a mid-range price, making it a new option. Haiku is extremely cheap and fast, and may unexpectedly be the most broadly useful model of the three.
Haiku in particular looks useful for work that needs to be handled quickly and inexpensively. Anthropic says: “It can read an information and data dense research paper on arXiv (~10k tokens) with charts and graphs in less than three seconds.” That is very appealing.
Its vision performance also looks very good. Claude has generally offered a better experience than ChatGPT or Gemini for work related to academic papers, and this may improve that experience further.
Harmlessness and the Context Window
Because of Anthropic’s emphasis on ethics and harmlessness, Claude has often frustrated users by responding that it cannot answer because of potential risks. Anthropic says this has been substantially improved.
Claude 3 starts with a 200k-token context window, with an implication that it may be expanded later. It also achieved very strong results on needle-in-a-haystack testing. These days, having a context window of several hundred thousand tokens without lost-in-the-middle problems is starting to feel like the baseline.
An Odd Needle-in-a-Haystack Response
An Anthropic researcher shared an interesting anecdote from internal 200k-context needle-in-a-haystack testing. In a test that asked about pizza toppings within a long document assembled from multiple pieces of writing, Claude 3 Opus reportedly answered with the correct response but also noted that the pizza discussion seemed out of place in the long text. It suggested that someone might be playing a trick or testing whether it was paying attention.
Needle-in-a-haystack benchmarks have become common as context windows grow. As LLMs become more capable, perhaps they will increasingly show emergent behavior by responding to benchmarks in ways that differ from ordinary tasks.
These benchmarks also contain an inherent awkwardness. If the needle is placed in natural writing, a model has a greater chance of guessing the answer from context rather than genuinely retrieving the information. If it is placed in unnatural writing, the test becomes detached from real-world scenarios—and now LLMs may notice that something about the setup feels suspicious.
Trying It Out
Opus is available immediately for $20 per month plus tax. For now, I plan to use the free Sonnet model for a while, then perhaps try the paid tier for about a month.