The shipped Create report screen: a "Describe the report you want to see…" box above three suggested questions, with saved reports and collections listed in the left sidebar

Rakuten Advertising • Jan 2025 - Ongoing

Intelligent Search for Custom Reports — Natural Language Search & AI

Users of Rakuten Advertising create custom reports on a regular basis, as often as weekly, to track campaign performance across 170+ metrics. Building one manually meant 15 to 20 minutes of clicking through dropdowns and configuring data points. With 1,000+ active users and dozens of account managers doing this regularly, the time loss was significant. It also landed on support when people couldn't figure out the interface.

Role: Sole UX designer

Skills: UX/UI, User Research, Prototyping, User testing

90%

faster report creation, measured in Fullstory during beta

~$10M

annual time-saving potential at full adoption

1,000+

active advertisers with access from open beta

Challenge

Natural language search features sound simple until you design one. The challenge wasn't just "add a text box", it was building trust in automation while preserving user control in an area where data accuracy matters. Users want to make decisions based on these reports, meaning any generated content needed to be verifiable and editable.
I needed to solve for:

  • Trust: How will a user know that the search results accurately match the query, without expecting them to check against up to 170+ metrics?
  • Ambiguity: How do we handle queries that are vague and could mean one of the many data points available?
  • Control vs. speed: Power users want a quick way of performing time consuming tasks, new users want guidance, how can we offer both?
  • Technical constraints: The query parser struggled with specific metric names that can be bespoke to Rakuten Advertising, or even the user's individual account, this required careful design around these considerations.
  • Fallback: Some users prefer manual control over search based automation. How can we ensure this option is still available to them?

The builder it replaced

The existing report builder started empty and stayed that way until you told it what to measure. Picking columns meant working through more than a dozen collapsed categories (Metrics, Clicks, Commission, Geography etc.) where some combinations quietly aren't permitted, and nothing rendered at all until at least one metric was chosen. Fifteen to twenty minutes later you had a table. And if the question you started with needed a number the table didn't directly answer, you exported it and worked it out in Excel.

That last step is the one that mattered: the tool produced data, not answers.

The legacy report builder: an empty New Report tab warning that at least one metric column is required, with the Add and Remove Columns panel open on a scrolling list of collapsed categories

Approach

I started by analysing existing reports to understand common patterns: What metrics did users combine? What date ranges mattered? What questions were they trying to answer? This informed the natural language query design. Instead of just free-form text, I included suggested questions to help a user get started and understand the expectations of the input box. I included a 'tag' system in a later iteration to help users find and include certain data points that were harder to remember.

Customer journey map comparing the current multi-step report creation flow with the proposed natural language search flow

Key decisions

  • Tags/Tokens: Users were able to include 'quick selected' tags to help direct a prompt better.
  • Manual override: Every generated report could be edited, saved, scheduled for a future date or rebuilt from scratch.
  • Feedback: Gathering feedback via Fullstory I was able to make further decisions in the UI and the functionality to help continually improve the feature.
  • Conservative defaults: The system suggested safe, common queries rather than trying to be clever.

I prototyped three ways of getting a question into the system: free text alone, free text with suggested questions to start from, and a structured tag system for naming specific metrics. Testing showed users reached for the suggestions first, as they helped a user understand the purpose and use case for the input box, while tags earned their place on the more complex requests, where remembering an exact metric name was the real barrier. Rather than pick one, the first design layered all three.

After being presented to our users, feedback was quickly gathered from internal teams and stakeholders in order to further steer the UI.

An early concept for the ask-a-question screen: a full-width purple page with a single 'Describe the report you want to see' box, three suggested questions as chips beneath it, and saved reports below

Solution

The solution combined natural language prompts, structured tags, and suggested queries to give users both speed and control. Every search generated report remained fully editable, could be saved as a template, or rebuilt from scratch, this preserved the manual workflow for users who preferred it.

Amends were included based on the feedback gathered after the initial release and an iterative approach meant I could deliver variations quickly and efficiently.

Tag and token system allowing users to refine and direct natural language search queries

Outcome

Closed beta launched in May 2025 with select power users, followed by a full open beta in July 2025 to all users. With this staggered approach it has allowed us to begin gathering adoption data and user feedback before the full release.

Based on the initial few months of usage we have determined that we have reduced the report creation time by up to 90% (measured using Fullstory during the beta phase) This translates to around $10 million in annual time saving potential when fully adopted by all users (both internal account managers and external users).

User feedback

"Super helpful to put in the prompts and get the reporting answers right away instead of having to sometimes pull a few different reports to get the answer." Account Manager
"When I needed to check week-on-week sales, Prompt made it easier and faster to generate the report, saving time and reducing manual effort." Account manager
"I was able to visualize best performing placement periods over time. I was able to add a 'lifetime value bounty' on top of RAD data. I was really impressed with that." Account manager

Learning from beta

What worked

  • Suggested prompts became a reliable onboarding tool, new users used them to understand what Prompt could do before digging deeper and building their own custom reports.
  • The tag system was adopted quickly by power users, enabling more precise, complex reports than pure natural language alone could produce.
  • Users began exploring reports they would never have built manually, discovery-driven reporting emerged as an unexpected use case.
  • Users who previously needed 3 to 4 report iterations were completing the same analysis in a single query.

What surprised me

  • Saving reports was used far less than expected, users found it easier to recreate reports on demand than to manage a saved library.
  • The addition of tags was not an initial plan, I assumed free-form text would be enough, but data-heavy reporting required a more precise input mechanism.
  • Trust was the real challenge for adoption, not usability. Users who didn't trust the output verified everything manually, which negated the time saving entirely. Adoption followed trust, not the other way around.

What I learned

The biggest surprise was how much trust mattered. I expected users to love the freedom of a text box, being allowed to describe a report rather than manually create it. What they actually needed was confidence that the output matched their intent. The full text-to-report approach had to evolve so users could see exactly what had been selected and step in if anything looked off.

I ended up with a hybrid interface. Everyone assumed users would prefer pure natural language, but the tag system became the most-used feature for complex reports. When there's a lot of data on the line, users want precision not just speed.