Science Blog

Send lawyers, guns and money

Navigation

  • Topics
    • Aerospace
    • Animals
    • Anthro and Archaeology
    • Bio and Medicine
    • Brain and Behavior
    • Business and Economy
    • Computers and Electronics
    • Education and Outreach
    • Energy and Environment
    • Geoscience
    • Internet and Communication
    • Media and Entertainment
    • Nanotech, Chem and Materials
    • Physics and Numbers
    • Security and Defense
    • Software
    • Space
    • Transportation
  • Reader Blogs
  • Shameless Commerce
  • Register/Login
Home Tool Helps Internet Master Top-level Domains
  • Contact
  • Home
Google

Recent Comments

  • Broken Heart
  • walang kwento
  • Where to publish without review?
  • RE: I haven't had a problem
  • ex felon no need to apply
more

Reader Blogs

  • Interference in Memory
  • Neuroscience that matters
  • 320,000 Acres of Forest Protected in Landmark Deal
  • Forgetting what you haven't yet learned
more

Tool Helps Internet Master Top-level Domains

Master of your domain

At the request of a worldwide Internet organization, a computer scientist at the National Institute of Standards and Technology (NIST) developed an algorithm that may guide applicants in proposing new “top-level domains”—the last part of an Internet address, such as .com, that people type in navigating the Web.

As new top-level domains are added to the familiar .com, .info and .net, the algorithm* checks whether the newly proposed name is confusingly similar to existing ones by looking for visual likenesses in its appearance. Having visually distinct top-level domain names may help avoid confusion in navigating the ever-expanding Internet and combat fraud, by reducing the potential to create malicious look-alikes: .C0M with a zero instead of .COM, for instance.

Later this year, the Internet Corporation for Assigned Names and Numbers (ICANN) plans to launch the process for proposing a new round of “generic” top-level domains (gTLDs), strings such as .net, .gov and .org meant to indicate organizations or interests. In preparing for newly proposed gTLDs, ICANN reached out to various algorithm developers, including NIST’s Paul E. Black, as among those engaged to “provide an open, objective, and predictable mechanism for assessing the degree of visual confusion” in gTLDs.

Black’s algorithm compares a proposed gTLD with other TLDs and generates a score based on their visual similarities. For example, the domain .C0M scores an 88 percent visual similarity with the familiar .COM. The resulting scores may help indicate whether the newly proposed domain name looks too much like existing ones.

To make its assessments, the algorithm rates the degree of similarity between pairs of alpha-numeric characters. Some pairs, such as the numeral “1” and its dead-ringer, the lowercase letter “l,” are assigned the highest scores for visual similarity while other pairs, such as “h” and “n”, are given lower scores. The algorithm takes other considerations into account, for example how certain pairs of letters, like “c” and “l,” can join to look like a third letter (“d”), as in the case of “close” and “dose.” Employing these scores and considerations, the algorithm computes the “cost” of transforming one string of characters into another, such as “opel” into “apple.” Lower cost means higher visual similarity. The algorithm then adjusts for the relative lengths of the two strings (different lengths increase their distinctiveness) and converts the final cost into a percent similarity.

ICANN is considering future enhancements to the algorithm, such as having it check for visual confusion between existing domains and future planned Internet top-level domain names in scripts such as Cyrillic.

* The algorithm can be found on the NIST Web page “Compute Visual Similarity of Top-Level Domains.”

Media Contact: Ben Stein, bstein@nist.gov, (301) 975-3097

http://www.nist.gov/public_affairs/techbeat/tb2008_0513.htm#domains



Reply

  • Web page addresses and e-mail addresses turn into links automatically.
  • Allowed HTML tags: <a> <em> <strong> <cite> <code> <ul> <ol> <li> <dl> <dt> <dd> <img> <blockquote>
  • Lines and paragraphs break automatically.

More information about formatting options

Copyright, Science Blog.
Think. It's not illegal yet. Read our Privacy Policy.
RoopleTheme