Note: This was originally posted by an inactive account. Content was preserved by moving under an admin account.
Originally posted by: stonysmithUsing the fuzzy match node is likely not going to work for you. At a depth of 4 characters, you are going to get significant multiple matches such as: "Main" and "Polk" streets match each other at a depth of 4 character changes.
What you are more likely going to have to do is parse the address into pieces:
House Number (1000)
House Fraction (A,B, 1/2)
PreDirection (NE)
Street Name (Broadway)
Street Type (St, Ave, Dr, Blvd)
PostDirection (SW)
Dwelling Type (Apt, Ste, Lot)
Apartment # (312)
Then, you have to standardize the various fields individually. (AVE instead of Avenue, ST instead of Street, etc)
Once you have the address parsed properly, you can try to find mis-spelled street names, etc.
Attached is a graph that I started constructing.. it's got a good way to go yet.
Attachments:
Address Parsing.brg