|
May 31 and June 1, 2007
8:30am to 4:00pm
Rutgers University/New Brunswick – Busch Campus
Getting early
warning of the outbreak of a disease or a biological attack -
Donna Schneider
Slides:
Mathematical
Prerequisites
Knowledge of functions and graphs of functions
and familiarity with some basic ideas in descriptive and inferential statistics.
Supplementary Information
Statistical Approaches to Authorship
Attribution -
David Madigan
Slides: Statistical Approaches to Authorship Attribution
Mathematical
Prerequisites
To understand Bayes Rule you would need
to do the chapter on Probability Theory in any Elementary Statistics textbook. The chapter usually starts with the
definition of probability in terms of sets, so some basic set theory notation is
essential. Then we do the rules of probability like the sum rule.
Following that is conditional probability and then comes Bayes Rule.
You also need to have an understanding of random
variables and in particular discrete random variables. You could do this
from the purely probabilistic perspective, but to really understand this you
should know some statistics - basic terminology, representing data by bar graphs
and histograms, measures of central tendency, measures of dispersion, and a good
understanding of "randomness."
The above material forms the bulk of descriptive
statistics (excluding correlation and least-squares regression) in a college freshman level statistics course or an AP Stats course
and usually gets done half way through the semester (or year).
Naive Bayes teachnique simply means we assume the
events are independent. Besides authorship attribution, Naive Bayes techniques are also used in,
for example, medical expert systems including systems that produce differential diagnoses based
on a list of symptoms. A differential diagnosis is a list of the diseases that
could possibly cause a patient's symptoms, arranged in decreasing order by
probability. For example, for a patient with a runny nose and sore throat,
common cold would be at or near the top of the list.
Supplementary Information
-
The Characteristic Curves of Compositions
by Thomas Mendenhall appears in Science - an illustrated journal, Volume
IX Jan - June 1887. This is the first paper on the subject.
Mendenhall compared 5 thousand word passages from Oliver Twist and
counted how many one letter words etc and found a definite pattern.
-
Mark Twain and
the Quintus Curtius Snodgrass Letters: A statistical test of authorship
(this paper concluded that he did not write these letters)
-
Some
discussion on the chi-square test used in the above paper
-
The Federalist
-
Inference in an authorship problem - this paper
concluded that Madison is the principle author
-
DIMACS working group on
Authorship Attribution - identifying real-life authors in massive document
collections
-
June 2006 Workshop
by the DIMACS working group on
Developing Community Resources for Automated Authorship Attribution
-
Proposal - gives background information on the subject
-
Authorship Attribution Bibliography - several papers are available online
-
Authorship
Attribution of Texts: a review
-
The state of
authorship attribution studies
-
The automatic detection of plagiarism - an undergraduate dissertation
-
n-gram based author
profiles for authorship attribution
An n-gram is a fixed-width substring of a
document. As an alternative to counting occurrences of words in a document, the
n-grams at offset 1,2,3, etc. of a given width can be counted. For instance, the
5-grams from the sentence "I LIKE MY CAR" are "I LIK", " LIKE", "IKE
M", "KE MY", "E MY ", " MY C", "MY CA" and "Y CAR". This technique has the
benefit that it largely replaces word stemming, which involves resolving
different word forms, like "drive" and "driving" to the same root. It also takes
into account to a limited extent the transition from one word to the next.
Smaller width n-grams, like 2-grams, are sometimes used to help resolve spelling
errors.
-
n-grams from Wikipedia
Inspecting containers at ports - Fred Roberts
Slides: Algorithms for Port of Entry Inspection for WMDs
Graph Theoretical Problems arising from defending against
bioterrorism and controlling the spread of fires -
Fred Roberts
Slides: Graph-theoretical Problems Arising from Defending Against Bioterrorism and Controlling the Spread of Fires
Locating Sensors to detect chemical, biological, or nuclear
threats - Fred Roberts
Slides: Locating Sensors to Detect Chem, Bio, or Nuclear Threats
Mathematical Prerequisites
The first half of the lectures require familiarity with elementary rules of
counting including the product rule, the sum rule, permutations and
combinations, bit strings, boolean functions and graph theory concepts including
paths, trees, and binary rooted trees. The second half of the lectures
need more advanced topics such as cost functions, sensitivity analysis including
the Stroud-Saeger method, network theory, theory of algorithms, and location
theory.
Supplementary Information
Relationship
Discovery in Large Text Collections Using Latent Semantic Indexing -
Bill Pottenger
Slides:
Mathematical
Prerequisites
Latent semantic
indexing is based on a matrix factoring method called singular value
decomposition (SVD). This topic is usually covered toward the end of a
college-level sophomore linear algebra course. To understand this
factoring technique you must first know LU factoring and the computational edge
it has over Gaussian Elimination. These topics fall under the broad area
of numerical linear algebra. So it would not be unusual to cover them in a
college-level junior or senior numerical analysis course.
Supplementary Information
Sponsored by:
- Department of Homeland Security Center for Dynamic Data Analysis (DyDAn) at Rutgers University
- Center for Discrete Mathematics and Theoretical Computer Science (DIMACS)
- Rutgers Center for Mathematics, Science, and Computer Education
(NJ Professional Development Provider #2)
|