Evaluating Legal AI for Philippine Tax Research

Tyra Delos Reyes

Philippine tax research has a habit of sending lawyers everywhere at once. What starts as a BIR question can quickly pull in SEC rules, CTA decisions, BOC issuances, DOF opinions, Supreme Court cases, and tax treaties.

Our community of CPAs and tax counsels know this rabbit hole well. One tax practitioner shared that a single issue can surface more than a hundred documents, all of which must be reviewed and narrowed down to the few authorities that actually matter. Miss one key issuance, and the entire analysis may rest on the wrong legal foundation.

This reflects how many tax lawyers use Anycase: to build defensible legal positions by synthesizing authorities across multiple agencies. 

While the evaluation included tax computation questions, its primary focus was synthesis and legal relevance. We focused on these legal AI competencies because these are the use cases that tax professionals on Anycase cite most often, and where they encounter the greatest friction in practice. 

How we evaluated Anycase and GPT

We evaluated Anycase against GPT-5.5 across 60 Philippine tax-related queries. The evaluation, designed by lawyers from our Legal Intelligence Team, included multiple-choice and essay-type questions involving five key government bodies:

  • Bureau of Internal Revenue

  • Securities and Exchange Commission

  • Court of Tax Appeals

  • Bureau of Customs

  • Department of Finance

The score measures how each AI tool handled specific tax problems, including queries around:

  • VAT treatment of marketplace sales

  • Dissolution procedure of a corporation with unpaid taxes

  • Appeal of a customs classification dispute

  • The CTA’s jurisdiction over tax and customs cases

  • DOF process for tax-exemption applications

Anycase turns fragmented tax authorities into legally grounded work products 

Across the full evaluation, Anycase returned an average score of 89.26%, compared to 59.03% for GPT 5.5.

That gap demonstrates that confident-sounding answers, notoriously known in LLMs, are not reflective of legal soundness. 

In a heavily-regulated field like tax, an AI’s legal analysis is only as strong as the authority behind it. Without the controlling legal basis, there is little foundation for a legal opinion. 

Across the five agencies, two recurring failure modes in GPT’s responses explained much of the performance gap: 

1) Citing non-existent, hallucinated references

GPT relied on nonexistent or hallucinated legal references in 15% of the queries.

This is more than a research inconvenience: a work product that rests on an invented authority cannot be verified, defended before a BIR inquiry, or safely used to navigate a customs investigation. 

Anycase, by comparison, more consistently cited the specific BIR or DOF issuance governing the issue, together with the relevant provisions of the National Internal Revenue Code. 

This gives practitioners a verifiable legal basis from which they could move beyond research and determine the appropriate next step, whether drafting an opinion or determining the correct forum and mode of appeal. 

2) Citing provisions that were related, but non-controlling authorities

Even when GPT cited genuine legal materials, it did not always identify the authority that governed the specific issue. 

This incorrect application of law or facts appeared in 21.7% of GPT’s 21.7% of the GPT’s responses, compared to Anycase’s 1.67% for Anycase. Anycase produced an incorrect answer in just one of the 60 questions because it was unable to retrieve the Revised Rules of the Court of Tax Appeals at the time, relying on other canonical sources to resolve the query. That retrieval gap has since been fixed. 

In these instances, GPT often relied on provisions that were related to the subject but too broad, irrelevant, or otherwise non-controlling. In some cases, it drew from the wrong agency altogether, such as applying a BOC issuance to a matter governed by CTA Rules. 

These two failure modes led to the same practical risk: the answer may look legally supported while resting on an authority that cannot withstand scrutiny. 

Anycase’s stronger performance reflects its investment in systems that augment AI retrieval with lawyer judgment. 

Through LAGDA, or Lawyer-Annotated and Governed Data Architecture, Anycase’s Legal Intelligence Team annotates and prioritizes controlling authorities. This helps the system distinguish between sources that merely mention the issue and those that directly resolve it. 

Performance by agency 

As Anycase’s Philippine legal library grows, our users should not have to trade breadth for noise. 

We evaluate AI tools at the agency level to test whether our legal AI can move across overlapping authorities (i.e. consult both BIR and SEC rules) and produce an answer that accounts for the operative requirements of each, rather than treating the issue as belonging to only one agency. 

Here’s how the two tools fared: 

Bureau of Internal Revenue and the Securities and Exchanges Commission

We grouped BIR and SEC queries to reflect real corporate tax transactions, where a single restructuring or dissolution may require applying the National Internal Revenue Code, alongside BIR revenue regulations, SEC rulings, and the Revised Corporation Code.

Anycase was generally stronger at identifying the correct answer, citing the specific issuances behind it, and explaining how the requirements of both agencies operated together. 

GPT sometimes reached the right conclusion, but often relied on incomplete, nonexistent, or misapplied authorities. In one example, it cited a general provision of the Revised Corporation Code instead of the BIR issuance that directly governed the tax consequence. 

Department of Finance

That pattern became even more pronounced in queries involving documents from the Department of Finance, where Anycase scored 98.80% compared with GPT’s 60.3%. 

GPT often defaulted to broad statutory language rather than the specific DOF issuance that supplies the controlling administrative interpretation. In practice, the answer may sound legally defensible, yet misstate the operative rule on coverage, eligibility, or compliance, driving avoidable tax exposure or unnecessary conservatism when the issue turns on how the agency applies the law.  

As a result, its answers offered limited practical value, particularly for procedural questions where lawyers need precise, actionable guidance. 

Court of Tax Appeals

Queries involving documents from the Court of Tax Appeals showed the consequences of that weakness at the case level. 

Although GPT answered 50% of the multiple-choice questions correctly, it often failed to support those conclusions with the authority that actually governed the legal issue. It cited irrelevant or misapplied cases, including a Bureau of Customs issuance that did not resolve the query. 

Anycase, by contrast, more consistently surfaced the correct authorities and substantive CTA cases. This enables lawyers to see not just the “right answer,” but the actual precedent and how the rule is applied in a ruling by the CTA. 

Bureau of Customs

The Bureau of Customs results reinforced the same pattern seen in CTA evaluations. GPT often cited legal authorities but misread the facts, mismatched the applicable rule, or misstated the issuance itself.

In comparison, Anycase was more reliable at locating that authority, giving practitioners a clearer path through various matters, from the examination of goods to investigations around misclassification and undervaluation of shipments. 

Evaluating Anycase and GPT against tax calculations

Anycase and GPT failed in different ways on the word-based tax problems. Anycase more often identified the correct tax framework, but its computations broke down when applying graduated tax rates, selecting the correct VAT base, accounting for available tax methods, and carrying exemptions or deductions into the final taxable income. In some answers, its summary also conflicted with its own computation. 

GPT’s failures were more conceptual: it misunderstood the treatment of final taxes, added final-tax liabilities to the regular income tax payable, failed to exclude non-taxable bonuses and benefits, and incorrectly applied rules for mixed-income earners. 

In practice, Anycase’s errors were more often found in the mechanics of applying the rule, while GPT more often misclassified the income or applied the wrong tax treatment altogether.

Anycase’s next steps for reliable tax calculations

Mathematical reasoning is a known weakness for both general-purpose LLMs and specialized legal AI tools that lack dedicated computational systems, because solving word problems requires several capabilities to work correctly at the same time.

For Anycase, improving tax calculations would therefore require a dedicated system for computational queries, similar in ambition to its Superseded Law Handling system and Lawyer-Annotated and Governed Data Architecture (LAGDA). 

Research showing that trained verifiers improve performance on mathematical word problems supports this direction. 

Until that layer is developed and tested, Anycase’s stronger use case remains procedural and case law tax research: identifying the governing authorities and providing actionable next steps.



Level 21, 8 Rockwell, Hidalgo Dr., Rockwell Center, Makati City, Metro Manila, Philippines

Level 21, 8 Rockwell, Hidalgo Dr., Rockwell Center, Makati City, Metro Manila, Philippines

Level 21, 8 Rockwell, Hidalgo Dr., Rockwell Center, Makati City, Metro Manila, Philippines

Level 21, 8 Rockwell, Hidalgo Dr., Rockwell Center, Makati City, Metro Manila, Philippines