InQuira’s and Mercado’s approaches to structured search
InQuira and Mercado both have broadened their marketing pitches beyond their traditional specialties of structured search for e-commerce. Even so, it’s well worth talking about those search technologies, which offer features and precision that you just don’t get from generic search engines. There’s a lot going on in these rather cool products.
In broad outline, Mercado and InQuira each combine three basic search approaches:
- Generic text indexing.
- Augmentation via an ontology.
- A rules engine that helps the site owner determine which results and responses are shown under various circumstances.
Of the two, InQuira seems to have the more sophisticated ontology. Indeed, the not-wholly-absurd claim is that InQuira does natural-language processing (NLP). Both vendors incorporate user information in deciding which search results to show, in ways that may be harbingers of what generic search engines like Google and Yahoo will do down the road. Read more
Categories: InQuira, Mercado, Natural language processing (NLP), Ontologies, Search engines, Structured search | 2 Comments |
Is DMOZ the cure to Wikipedia’s spam problem?
Joost de Valk makes an interesting suggestion, namely that Wikipedia should drop all external links other than to DMOZ, and rely on DMOZ as the outside link directory. As division of labor, it makes perfect sense. However, it’s a total non-starter until at least two problems are solved. Read more
Categories: Categorization and filtering, Directories, ODP and DMOZ, Ontologies, Spam and antispam | 5 Comments |
Does anybody actually use Technorati?
I just did some Technorati searches, and my blog posts come up near the top of the search results for a bunch of small companies’ names and similar words — Attensity, ClearForest, Netezza, DATAllegro, Crossbeam, DMOZ, ODP, and surely many others.
But judging by my referrer logs, nobody cares. I get lots of visitors via classic search engines — largely Google, but also the others — but bubkus from Technorati.
Technorati Tags: Technorati
Categories: Search engines, Specialized search | 4 Comments |
Social networking architecture of the future continued
Responding to a question by Jon Udell a few hours ago, I argued that private social networking “walled gardens” aren’t needed. The whole thing can be done publicly as well, assuming there’s a central database to help with things like access control, as in the hypothetical service I named “Linkerati.”
Some other comments on his post raise issues like “Yes, but what if a walled garden is the easiest way to get people to post the needed information?” I have a quick reply: Just let all the needed information be entered in the central database, and you’re clearly better off than in a walled garden. Read more
Categories: Social software and online media | Leave a Comment |
Fact and Fiction: DMOZ and the ODP
- DMOZ is dead. Fiction!
- New site submissions are being processed. Partial fact.
- Pending site submissions were lost in the outage. Partial fact.
- Other non-public ODP data was lost in the outage too. Partial fact.
- New editor applications aren’t being processed yet. Fact.
- ODP editors are corrupt. Fiction!
- The ODP is secretive and deceptive. Largely fiction.
- If a DMOZ category doesn’t have a listed editor, it’s unlikely to get much attention. Part fact, part fiction.
- ODP editors hate search engine optimization. Partial fact.
- ODP editors hate SEOs. Partial fact.
I shall explain. Read more
Categories: Categorization and filtering, Directories, ODP and DMOZ, Search engine optimization (SEO) | 7 Comments |
A hobbit writes from the ODP Entmoot
Before saying anything about the Open Directory Project or the DMOZ directory it produces, I should offer several disclaimers.
- No editor speaks for the ODP, let alone for Time Warner/AOL/Netscape.
- No single editor’s opinions or choices control any edits in DMOZ, even if s/he is the sole listed editor of a category. Any of us can be overruled on any editing decision at any time.
- I’m effectively as new as they come, or at least was at the time DMOZ editing came back online (late December). There have been no new editors since the well-publicized outage, and I had next to no involvement with the project prior to the outage.
- Notwithstanding point #2, I’m quite opinionated, which I’m sure surprises approximately nobody. And my opinions quite often are different from those of the ODP mainstream.
Categories: Categorization and filtering, ODP and DMOZ | 1 Comment |
What is LinkedIn needed for? Absolutely nothing. And the same goes for MySpace.
Jon Udell asks whether private social networks such as LinkedIn are needed, or whether they can be completely refactored across the public internet. I say the latter. In social networking as in almost everything else, there’s no long-term need for an internet walled garden.
Categories: Social software and online media | 2 Comments |
Can Hakia hack it?
Hakia purports to be a new search engine that indexes “semantically,” which I presume means on phrases or concepts or something. But I’ve run a few queries side by side on Hakia and Google, and they’re not doing well. I think they’re not making sufficiently good use of page reputation. Try “web hosting forum” for an example of this, looking at the top two hits in both cases.
When I queried on “Viagra,” Hakia did — as it were — outperform Google. But that’s the only case I, uh, came up with. On less snigger-worthy searches, Google seemed to do as well as or better than Hakia.
Categories: Google, Search engines | Comments Off on Can Hakia hack it? |
Please switch to my back-up e-mail address
At least for the moment.
monash.com e-mail has been turned off by my hosting company, due to what they claim is a still on-going attack. My backup address, however — FirstnameLastname@domain.com, where domain = dbms2 — is working fine. And my e-mail client traditionally checks them at the same time. So I suggest switching, at least for the moment.
Both are through the same hosting company (Hostgator, which I aspire to replace in the immediate future, given that I also lost admin access to the blogs on two separate occasions this week, and given that support claims over half my e-mails are unreadably empty and hence suitable for being ignored, despite me never having that problem elsewhere). Thus, for other kinds of problems there might be a single point of failure. But in this case, the dbms2 address is a working alternative to the standard one.
Categories: About this blog, Spam and antispam | Leave a Comment |
What’s interesting about the FAST venture in BI
FAST is annoying me a bit these days. It’s nothing serious, but travel schedule screw-up’s, an annoying embargo, and a screw-up in the annoying embargo have all hit at once. So I’ll keep this telegraphic and move on to other subjects.
- They’re doing fast queries without using a lot of RAM.
- They’re doing the usual text search thing of indexing across multiple “databases,” only now it’s applied to, well, databases. (Not that there’s much new about that particular aspect. Actually, there seems to be a bit of kludge in that they export the databases to some kind of simple text files.)
- They’re doing some level of concept identification ala the text mining guys. (They don’t call it “entity extraction” because the results aren’t dumped into a database anywhere, but instead are just used on the fly.) Of course, the text mining/search convergence goes both ways.
- They bought a BI/dashboard tool and are using it both to analyze query logs and also to do normal BI/dashboard kinds of things.
- They have big references for this stuff, at least the single-web-site query aspect. Well, actually, the customer names are confidential. Oh well.
And as another example of how this wasn’t the smoothest PR month for FAST, Steve Arnold somehow got the false idea that they were getting out of true text search altogether.
Categories: BI integration, Enterprise search, FAST, Search engines, Text mining | 3 Comments |