Newsroom
Introducing L3
The Arabic language engine behind WriteX now has a published evaluation: first on every primary metric against 26 large language model configurations, 80.29% edit F0.5 against 44.68% for the strongest of them, with the method and its limits in the open.
- L3, the Lisan language engine inside WriteX, ranked first on every primary metric against 26 large language model configurations on a linguist-reviewed benchmark of real Arabic sentences. Measured.
- 80.29% edit F0.5 against 44.68% for Claude Fable 5.1, the strongest configuration; the only system with precision above 80% and recall above 60% together.
- Not independently verified. Second on throughput, behind Qwen 3.8 27B on Groq. The energy anchors are cited estimates. Prompt, scoring definitions and limits published in full.
L3, the Lisan language engine that checks Arabic inside WriteX, ranked first on every primary metric against 26 large language model configurations on a linguist-reviewed benchmark of real Arabic sentences. On edit F0.5, the headline measure, it scored 80.29%, against 44.68% for Claude Fable 5.1, the strongest configuration in the field. Measured.
Today we are publishing the full evaluation report. The field is 26 configurations from Anthropic, OpenAI, Google, DeepSeek, Meta and Alibaba Qwen. Every system received the same sentences, the same strict Arabic prompt, one output per sentence and the same scoring. The evaluation snapshot is 9 September 2026, version 1.0.
F0.5 combines precision and recall and weights precision more heavily. Precision is the share of proposed changes that landed where the linguist also changed the text; recall is the share of the linguist's edits the system found. Full definitions in the report.
L3 leads on edit F0.5 at 80.29%, against 44.68% for Claude Fable 5.1
Edit F0.5, exact-span detection, top ten configurations. Higher is better. F0.5 weights precision above recall, because unnecessary changes cost more than missed ones.
Why the shape of the result matters
Language checking is two tasks: finding the span that is wrong, then writing the replacement. We score them separately, and we weight precision above recall, because an unnecessary change in an official document costs more than a missed one. L3 was the only system that combined precision above 80% with recall above 60%. The large language models split the other way. Claude Fable 5.1 was the most precise of them at 64.21% but found only 20.15% of the reference edits; GPT-5.6 Sol found the most, at 41.08%, with 25.03% precision.
Where L3 located a span correctly, its replacement matched the linguist's reference 89.79% of the time exactly and 96.38% after Arabic normalization. Those rates are conditional on detection, and the report says so at every mention. On the strict whole-sentence test, where one differing character fails the sentence, L3 matched the reference on 55.96% of sentences, double the rate of Claude Fable 5.1.
One sentence from the benchmark
A single row shows what the numbers describe. The source asks what predictive analytics means, and spells the adjective with the hamza seated as it is in the noun: ما المقصود بالتحليلات التنبؤية؟. The classical rule moves the hamza to a yeh carrier in the adjective, and the linguist's reference reads ما المقصود بالتحليلات التنبئية؟. L3 returned the reference exactly. No large language model configuration produced the reference spelling. The source spelling is very common in modern technical Arabic, which is the point: an institutional standard is a set of rules held consistently, and a language checker is judged on holding them.
The report also shows rows where the models did well, rows where the reference applies an editorial convention rather than correcting an error, and rows where L3 missed. We labelled each of them as what it is.
Speed and energy
L3 processed 52.28 standardized output tokens per second, second in the field to Qwen 3.8 27B on Groq at 66.98. Measured directly on an NVIDIA A10, it used approximately 0.000014 kWh per 1,000 tokens at accelerator scope, or 0.000042 kWh under a three-times full-stack estimate. Against cited large language model anchors, from Google's Gemini Apps disclosure to the IFP School estimates for GPT-5, that is approximately 29 to 3,357 times lower. The anchors are estimates, not measurements of the tested models, and the report shows the sensitivity range. Measured for L3; cited estimates for the anchors.
What we are not claiming
The evaluation was designed and run by Lisan Research and has not been independently verified. It is one snapshot, on one benchmark, with one linguist reference per sentence. The prompt forbade rephrasing, so rows that apply a house editorial standard favour the engine that implements that standard. L3 missed reference edits and made changes the linguist did not. All of this is in the report, with every rate in the appendices and the verbatim prompt, so that anyone with a linguist-reviewed Arabic set can repeat the protocol.
Read the full evaluation report, see what the results mean inside the product on the WriteX results page, or check your own text in the free editor. Institutions that want L3 on their own infrastructure can contact us about a pilot.
Start with one product.
Keep the whole platform.
Open a free account and solve today's problem in the next ten minutes. When you are ready for more, six flagships and a 20+ app workspace are already on your account. Or talk to us and we will scope it with you.
Free to start, no card · Your data exportable, always · Trusted by more than 24 government entities