Alibaba has open-sourced Open Code Review, an AI-powered command-line tool for code review. Open Code Review reads Git diffs, sends changed files to a configurable language model through an agent that can use tools, and produces structured line-level comments.
Deterministic logic handles file selection, grouping related files, and rule matching. The agent makes decisions during analysis and retrieves the context it needs.
In Alibaba’s own comparison with Claude Code using the same model, Open Code Review achieved higher Precision and F1 while consuming about one ninth of the tokens. Its Recall is lower than that of general-purpose agents, which the developers describe as a trade-off for precision and fewer false alarms.
Claim check:
- Alibaba has open-sourced Open Code Review, an AI-powered command-line tool for code review. (confirmed by the publication itself: evidence; «Open Code Review is an AI-powered code review CLI tool. It originated as Alibaba Group’s internal official AI code review assistant — over the past two years, it has served tens of thousands of developers and identified millions of code defects. After thorough validation at massive scale, we incubated it into an open source project for the community.»)
- Open Code Review reads Git diffs, sends changed files to a configurable language model through an agent that can use tools, and produces structured line-level comments. (confirmed by the publication itself: evidence; «It reads Git diffs, sends changed files to a configurable LLM via an agent with tool-use capabilities, and generates structured review comments with line-level precision.»)
- Deterministic logic handles file selection, grouping related files, and rule matching. (confirmed by the publication itself: evidence; «For review steps that must not go wrong , engineering logic — not the language model — guarantees correctness:»)
- The agent makes decisions during analysis and retrieves the context it needs. (confirmed by the publication itself: evidence; «The agent’s strengths are concentrated where they matter most — dynamic decisions and dynamic context retrieval:»)
- In Alibaba’s own comparison with Claude Code using the same model, Open Code Review achieved higher Precision and F1 while consuming about one ninth of the tokens. (confirmed by the publication itself: evidence; «Compared to general-purpose agents (Claude Code), Open Code Review achieves significantly higher Precision and F1 with the same underlying model, while consuming only ~1/9 of the tokens and completing reviews faster.»)
- Its Recall is lower than that of general-purpose agents, which the developers describe as a trade-off for precision and fewer false alarms. (confirmed by the publication itself: evidence; «Note that its Recall is lower than general-purpose agents — a deliberate trade-off favoring precision over noise.»)
Publications:
Primary sources:
score 68.2 out of 100 · kind: announcement