Showing posts with label Half baked ideas. Show all posts
Showing posts with label Half baked ideas. Show all posts

Friday, March 06, 2009

Emails Are Forever

I don't know about you but the way I access my emails is almost always through the Gmail search box, whether I'm looking for that email about someone's flashy new job (since I forgot where they are working) or my todo list from last week.  And that's primarily because of the amount and nature of data in my Gmail.  All the communication of any consequence is reflected in emails (including facebook, orkut scraps), and additionally I've moved a lot of note taking to Gmail as well - implemented by sending an email to myself.

With so much of my life depending on my online existence, I'm getting a bit paranoid about several what if scenarios, like what if the online servicing maintaining my data goes bankrupt, etc.  Of course my case is not unique.  Millions of people worldwide have a lot of precious data sitting on computers spread out across the world, owned by a foreign corporation, and supported free of charge (e.g., free Gmail is partially supported by Google's online ads revenue).  But assuming some of this data will remain as valuable to me a decade or two later, I'd like to ensure that I don't leave all of it to the vagaries of a ticker symbol.

Maybe you would argue that Google and Yahoo and Microsoft are not going away anytime soon and I believe you.  But less than 20% of firms that existed two decades back are still around, so that increases the odds against your hypothesis.  Another possible argument is that a product getting phased out doesn't necessarily imply that user data will be lost (e.g., Yahoo photos to Flickr), maybe user data will become all the more precious over time.  Still, I'd argue that the expected lifetime of your present online data will only go down over time.  So the big question is do you care enough about _all_ that data or not.  In my case, I care about most of it.

A poor man's solution is to immortalize your data is to periodically "download" and archive it at your home or better still at another online service.  Of course, switching services may affect the usability of the data, which affects its value to you.  E.g., if I zipped all my Gmail into one single archive, I can't search freely like before.  It may still be useful for litigation purposes but not for looking up my roommate's phone number.  This brings us to the first rule:

1.  When archiving your data, the functionality supported by the archived data should be comparable to the primary copy's functionality.  If not, it will probably be forgotten.

A more difficult problem for me is figuring out _where_ to archive or move my data, making the decision, verifying that it worked, etc.  This is potentially very time consuming.  So what I'd want to see is a computer "program" that will keep track of the service level trends for your current data provider, monitors/evaluates upcoming alternative services, and replicates your data across a "diverse portfolio" of service providers, so over time your data is likely to survive somewhere - and almost always it can be found on the new and upcoming service.

It may seem that I just borrowed a dialog from Star Trek and I can understand your sentiment.  Some key components that will be needed to create such a "program" include:
  • A service for evaluating and recommending data service providers.
  • A data API supported by different data service providers that enables easy migration between them.  For example, if you must migrate a photo sharing service to an email provider, you'd probably store a single photo album as a single email.  Someone will have to write such interfaces for all new services.
Before I conclude, I'd like to observe that this problem is similar to money management.  When you give your money to a money management firm, you expect them to keep reinvesting it in the most appropriate business presently.  Maybe that is too complex to delegate fully to a program, but I think keeping your data up and running forever should be simpler.

Saturday, July 05, 2008

NLP based search

I am getting increasingly interested in Natural Language Processing (NLP) these days. NLP can enable better human computer interfaces, powerful search engines, etc. One of the search startups in this area that I have been following is www.powerset.com which was recently acquired by Microsoft. A good source to learn about powerset and a rought technical overview is at http://www.slate.com/id/2193837/.

Powerset's NLP technology breaks a sentence into smaller entities (nouns, verbs, adjectives, etc.) and establishes relationships between them, e.g., "eiffel tower was built in 1889" gets recorded as "eiffel tower" (noun) "built" (verb) 1889 (noun). Each such relationship (called "fact") is recorded and comprises a single quantum of information derived from the web page. A search query is translated into a similar, but incomplete fact, e.g., "when was eiffel tower constructed?" would become "eiffel tower" (noun), "constructed" (verb), and "when/year/time/date" (noun/adjective). The search algorithm then matches the "factualized" query to the closest resembling fact and fills in the missing details (the year 1889 in this case).

The cool thing about converting content and queries to facts is that the search engine can identify and return relationships not explicitly stated in the contents, unlike keyword based search. However, most popular content on the web is actually explicitly stated in a single sentence, so NLP seems less useful for searching popular content since Google search would already do a pretty good job here.

However, the real promise of NLP based search seems to be in the context of the "long tail" of search - which are frequently searches not explicity answered on any single web page. As the web continues to grow and many different kinds of contents come online (blogs, books, emails, etc.), the long tail of web search will continue to increase its share of the total search volume. Most of us have experienced that the unpopular searches often are not explicitly found in any single web page, instead they require the user to scan multiple web pages before they find what they want. Keyword based searching cannot make things any better here since the keywords may either be spread out across webpages or they may simply be absent (e.g., "dog" and "tommy" can be related if tommy is the name of a dog - a fact that keyword based search cannot discover). This is where NLP can really make a difference. It can identify facts from across web pages and save users valuable time spent scanning different web pages trying to forge an answer to their search queries.

So, very roughly speaking, if you can find the answer to your query in one Google search and after scanning 1-2 returned web pages, then NLP will not make things any better for you. If it takes more than one search and visiting 5+ search results to answer a given query, and if your query and its potential response can be formed into a fact, then NLP might be useful.

Another analogy for the applicability of NLP may be the information density of a web page. NLP will be more useful finding content in web pages with low information density. By converting the text to facts, NLP is in a way converting "semantic compression" of the contents. "NLP compressed facts", owing to their increased information density, are better suited to answer user queries. These "low information density" web pages may be web pages with lower page ranks on Google. Other examples of low density content might be casual chat sessions, email threads, etc.

Unfortunately, in my experience Powerset doesn't seem to be doing a good job in identifying complex facts. They do a decent job at identifying obvious or simple facts but based on some examples I saw, not so well for complex facts. For examle, if you search for "who was the author of the godfather", you get the answer "Mario Puzo". But Google also fairly easily gives you the same answer when you search for "Godfather author" or "Godfather writer". But if you query, "how many years did Mario Puzo take to write the godfather", Powerset doesn't seem to offer any useful results.

Also, I wonder if their algorithm can really connect information from across different websites, different paragraphs in the same web page (should be there I think), etc.

I'd conclude that for NLP search to be really useful, it should target the long tail of searches - searches which individually are an insignificant part of the total search volume but put together comprise a major chunk. Powerset NLP search doesn't seem to be there yet and quite likely neither do other existing NLP based searche engines.

Thursday, July 19, 2007

Trust Issues With CellSwapper

CellSwapper is a new service (http://www.cellswapper.com) that allows one to get out of a cell phone contract they are stuck with, without having to pay the termination fee (only paying a much smaller service fee). This is a cool service and will certainly be found useful by many users stuck in their long contracts who are frustrated due to poor reception, etc. The owner of the contract (the "seller") gets out of the contract by transferring it to someone who wants a new contract (the "buyer"), the transfer being facilitated by CellSwapper. The seller gains by not having to pay the huge termination fee, and the buyer gains by not having to pay the activation fee on a new contract. However, I see one potential problem with the transfer process which can be exploited by malicious users.

This is how the contract transfer happens:
  1. Seller posts his contract details on CellSwapper.
  2. Buyer contacts seller through CellSwapper
  3. Seller pays service fee to CellSwapper ($14-$19) to get buyer details.
  4. Seller transfers contract to buyer through the cell phone service provider.
  5. Seller ships the SIM card and/or phone to the buyer.
There are issues with this scheme mentioned on the CellSwapper website like if step 4 fails due to the buyer failing credit check, the seller loses his service fee, but I don't really consider it "malicious" on the buyer's part as he didn't gain anything. However, the seller can behave maliciously if he transfers the contract to the buyer in step 4 but doesn't immediately send the SIM card to the buyer. He could abuse the minutes on the contract before dispatching the SIM/phone and not have to worry about it since the financial responsibility is transferred already. Of course, this is possible only if step 4 happens without the seller and the buyer physically meeting each other, say doing the transfer of contract on the phone.

I don't see how the buyer can possibly protect himself from this unless the contract transfer process requires a re-activation by the buyer in which case the seller will not be able to use minutes on the contract anymore once the transfer is over. Although I haven't ever tried it myself, this kind of arrangement looks highly unlikely to me. So in the absence of this scheme, the buyer can only trust that the seller doesn't abuse the contract in this way.