Key Strategies for Building a System: A Tag Platform
Background
This is the tag system that manages and describes supply (merchants, dishes, products) for search in the Meituan app. It gives product, operations and engineering teams tools to manage tags, and through them delivers three things to users: tag-based recall, tag-based filters, and tags shown on results.
I was the product manager for supply tags. My job was to research what users care about in tags across the top search scenarios, manage tag data in line with business needs and operating strategy, and use that data to build three capabilities: expanding recall with tags, user-facing filters, and tag display.
Whether a system succeeds depends mostly on what you build first, who you build it for, and how little it takes before people can use it. Below are the two strategies that mattered most when I pushed the tag system forward.
1. MVP strategy: find the highest-ROI piece first
Starting point: a big, stalled plan
The original plan was engineering-led: turn the display templates users see into composable resources and assemble the results page like a jigsaw. When I dug in, I found the project aimed mainly at rebuilding capabilities and paid little attention to how business users would actually use it. The scope was also huge. It stalled before the configuration module was designed or the existing logic was even mapped out.
Redefining the value
From the tag system’s point of view, I asked: componentizing the results page costs a lot of engineering time, so where is the highest value and ROI? What would matter most to users and to the key decision makers? My conclusion: the biggest value lay in taking on business requests and shipping them quickly and accurately, then validating them through experiments to deliver business gains. Other gains, such as managing legacy logic or making product and engineering more efficient, were secondary.
Picking the MVP from the backlog
With that in mind, I went through the backlog of business requests. Because the project’s goal was efficiency, I picked the requests that came up most often and had already proven their business value. They fell into three groups:
- Tag recall: for a given kind of search query, recall more supply that has a certain feature. Examples: gyms with “refund if the gym closes” protection, restaurants with parking, products that can arrive within 30 minutes. Offline tags or real-time tag APIs can expand recall, make recall more precise, or filter results in the retrieval and ranking layers.
- Filters: for some queries we can tell that users want to broaden, shift or narrow their search. Narrowing is by far the biggest need.
- Narrowing by category. When users search “hot pot”, many of them rephrase to narrow the search. The first screen only shows two or three restaurants, already shaped by the user’s history, so category variety is thin. Feeding category tags into filters helps users pick the right kind of result much faster.
- Narrowing by brand. This also paid off well. Users care a lot about brands in food and medicine, so we measured experiments separately by category and by query. It had real challenges. For products with a short shelf life, like beer and milk, local brands often sell best in each region, so the filter has to follow local recall results. Some categories differ sharply between north and south: for mooncakes, people might look for Cantonese, Suzhou, Beijing or fresh-meat styles, plus different fillings and packaging, so ranking has to use region and user features. The final design: operations configure as complete a candidate set as possible; algorithms rank it; weights can be adjusted or options hidden by region and query; and the filter only shows options that match current recall results.
- Tag display: show tags on result cards as badges, reasons to choose, or rankings, so users can judge and decide quickly from the list.
Build order
We started with tag recall, partly to learn the search pipeline and partly to fix search experience problems. Then we ran user research with the design team and built capabilities in order of business value: first operations tools and strategy for filters, then tag display.
Takeaway
A big rewrite easily burns through patience and resources while still “mapping the current state”. Start by asking who will benefit right away, work back to the smallest usable scope, and the system gets a chance to be used. Real usage then drives what to build next.
2. Layering strategy: manage tags with metadata
Tag data needs a metadata system. Different tags naturally suit different uses: some work for recall but should never be shown to users, and so on. Getting this layering right is essential.
The three kinds of MVP requests map to the three uses of tags. Each use sets a different quality bar, so every tag’s metadata should say which layers it can be used in. Here is how that looks in common Meituan search scenarios:
| Recall | Filter | Display | |
|---|---|---|---|
| Problem it solves | The query doesn’t name a merchant or dish; find supply that matches the intent | The user has clear constraints and wants to narrow results fast | Help users judge and decide quickly from the result list |
| Good scenarios | Broad or occasion queries like “group dinner”, “date night”, “kid-friendly”; long-tail queries that structured fields can’t handle | Categories with many results and clear decision factors: free delivery or “within 30 minutes” for delivery; price per person, distance, open now for restaurants | Badges, reasons to choose and rankings on result cards, such as “lots of repeat customers”, “trending” or “must-eat list” |
| Typical tags | Occasion tags mined from reviews (good for groups, quiet), dish and flavor tags, synonym and category mappings | Category, brand, price range, delivery fee, delivery time, open status, area | Reputation tags, sales and repeat purchases, rankings, service promises (on-time delivery, food safety) |
| Requirements | Semantic relevance is enough; moderate precision is fine because ranking filters later; the wider the coverage the better | Discrete values users understand at a glance; coverage must be complete, since a missing tag removes the merchant from results | Highest precision; must be explainable to merchants and users; copy must comply with ad rules, with no absolute claims like “best” or “No. 1” |
| Precision and recall | Recall comes first; 70–80% precision is enough, since ranking removes mistakes | Both need to be high, usually around 95%: a wrong tag lets in merchants that don’t qualify, a missing tag hides ones that do | Precision comes first, usually above 95%, and above 98% for sensitive tags; recall can be low, since showing nothing beats showing something wrong, backed by an ops tool |
| Bad fits | Internal business tags can feed strategy, but never as evidence of user intent | Tags users rarely select (like “new store”), tags with low coverage | Model-mined tags with unstable precision; internal tags like “subsidized merchant” or “high-commission merchant” |
Whether a tag can be used, and where, varies a lot. Some typical cases:
- Recall only, never displayed. “Good for group dinners” is mined from reviews by a model and is right about 70–80% of the time. Using it to recall relevant restaurants is fine, since ranking filters again. Printing “good for group dinners” on a card turns the wrong 20–30% into misleading claims, and merchants will complain.
- Internal only, never displayed. Business tags like “subsidized merchant” or “high-commission merchant” can feed recall and ranking strategy. Showing them would put the platform’s business strategy on the table.
- Fine to display, poor as a filter. “New store” works as a badge that draws clicks, but few users would ever tick “new stores only”, so as a filter it just takes up space.
- A filter that can’t launch until coverage is complete. If “free delivery” covers only half the merchants, users who tick it will find many free-delivery stores missing. Offering no such option is better.
Low recall needs an operations backstop. To keep display tags precise, the model threshold is strict, so some merchants who deserve a tag will be missed. For a merchant, missing a tag like “lots of repeat customers” means real lost traffic, and they will complain. So besides the model, the display layer needs a support tool for operations: when a merchant reports a missing or wrong tag, operations can verify it and put the tag live quickly, without waiting for the next model update.
So the metadata should record at least: which layers the tag is allowed in (recall / filter / display / internal only), its source (entered by merchants, labeled by operations, mined by models), precision and coverage, update frequency, owner, and compliance review status for display tags.
With this layering, a new request starts with a metadata lookup: is the tag good enough for this layer? If yes, configure it and ship. If not, you know whether to improve precision or coverage, and you don’t have to evaluate everything from scratch each time.
Looking back
Looking back, both strategies answer the same question: where should limited resources go first?
- Find the value before building capabilities. The stalled jigsaw system had a technical plan. What it lacked was an answer to “who benefits right away?” Pulling the goal back from rebuilding capabilities to delivering business requests fast gave the MVP a clear boundary: only the most frequent, proven requests in recall, filters and display.
- The order is itself a strategy. Recall came first to learn the pipeline and fix the experience; user research then set the priority of filters and display. Each step built the understanding and trust the next one needed.
- Turn “can this tag be used here?” into data. A tag’s value depends on whether it is used in the right place. Recording each tag’s allowed layers, precision, recall and source in metadata replaces gut feel and repeated reviews; where models fall short, operations tools fill the gap.
The same approach works beyond tag systems. Any data or platform product with many kinds of users and a legacy burden can start with the highest-ROI use case as its MVP, then turn “what can be used where” into metadata anyone can look up.
Comments
Comments load here, powered by GitHub Discussions.