Comment on [deleted]
faithcure@lemmy.world 1 week agoGreat questions, let me break down how the pipeline works:
Our first layer is deterministic — we pull core data from free APIs. Then multiple layers of cross-querying happen on the backend, and the results get merged before anything hits the site. After that, an AI agent runs research across the web (always with source attribution) and fills in the highest-confidence data.
On top of that, there’s a report feature so users can flag incorrect info. Honestly, I’d argue our data quality is already higher than most sites in this space because of this layered approach.
That said, the system is still in testing — nothing here is final, and it’ll settle into place and improve over time.
On the “5 sources” point: which piece of info came from where is actually indicated as subtext on the page. But I hear you that it could be surfaced more clearly.
On sourcing and licensing — we’re strict about this. We only use free APIs. For example, we don’t pull any data from MobyGames because their terms are restrictive; we only do read-only matching to check whether existing data lines up.
As for licensing our own page content (the Terraria wiki example) — that’s a really good catch, and it’s not something we’ve formalized yet. Adding it to the to-do list. Genuinely valuable feedback, thank you.
popcar2@piefed.ca 1 week ago
Okay this is definitely AI spam lol
faithcure@lemmy.world 1 week ago
Not spam. I’m just using translation for best communication. Thank you.