|
||||||||||
PREV CLASS NEXT CLASS | FRAMES NO FRAMES | |||||||||
SUMMARY: NESTED | FIELD | CONSTR | METHOD | DETAIL: FIELD | CONSTR | METHOD |
java.lang.Objectorg.apache.lucene.analysis.TokenStream
org.apache.lucene.analysis.Tokenizer
org.apache.lucene.wikipedia.analysis.WikipediaTokenizer
public class WikipediaTokenizer
Extension of StandardTokenizer that is aware of Wikipedia syntax. It is based off of the Wikipedia tutorial available at http://en.wikipedia.org/wiki/Wikipedia:Tutorial, but it may not be complete.
EXPERIMENTAL !!!!!!!!! NOTE: This Tokenizer is considered experimental and the grammar is subject to change in the trunk and in follow up releases.
Field Summary | |
---|---|
static int |
ACRONYM_ID
|
static int |
ALPHANUM_ID
|
static int |
APOSTROPHE_ID
|
static String |
BOLD
|
static int |
BOLD_ID
|
static String |
BOLD_ITALICS
|
static int |
BOLD_ITALICS_ID
|
static int |
BOTH
|
static String |
CATEGORY
|
static int |
CATEGORY_ID
|
static String |
CITATION
|
static int |
CITATION_ID
|
static int |
CJ_ID
|
static int |
COMPANY_ID
|
static int |
EMAIL_ID
|
static String |
EXTERNAL_LINK
|
static int |
EXTERNAL_LINK_ID
|
static String |
EXTERNAL_LINK_URL
|
static int |
EXTERNAL_LINK_URL_ID
|
static String |
HEADING
|
static int |
HEADING_ID
|
static int |
HOST_ID
|
static String |
INTERNAL_LINK
|
static int |
INTERNAL_LINK_ID
|
static String |
ITALICS
|
static int |
ITALICS_ID
|
static int |
NUM_ID
|
static String |
SUB_HEADING
|
static int |
SUB_HEADING_ID
|
static String[] |
TOKEN_TYPES
String token types that correspond to token type int constants |
static String[] |
tokenImage
Deprecated. Please use TOKEN_TYPES instead |
static int |
TOKENS_ONLY
|
static int |
UNTOKENIZED_ONLY
|
static int |
UNTOKENIZED_TOKEN_FLAG
This flag is used to indicate that the produced "Token" would, if TOKENS_ONLY was used, produce multiple tokens. |
Fields inherited from class org.apache.lucene.analysis.Tokenizer |
---|
input |
Constructor Summary | |
---|---|
WikipediaTokenizer(Reader input)
Creates a new instance of the WikipediaTokenizer . |
|
WikipediaTokenizer(Reader input,
int tokenOutput,
Set untokenizedTypes)
|
Method Summary | |
---|---|
Token |
next(Token reusableToken)
Returns the next token in the stream, or null at EOS. |
void |
reset()
Resets this stream to the beginning. |
void |
reset(Reader reader)
Expert: Reset the tokenizer to a new reader. |
Methods inherited from class org.apache.lucene.analysis.Tokenizer |
---|
close |
Methods inherited from class org.apache.lucene.analysis.TokenStream |
---|
next |
Methods inherited from class java.lang.Object |
---|
clone, equals, finalize, getClass, hashCode, notify, notifyAll, toString, wait, wait, wait |
Field Detail |
---|
public static final String INTERNAL_LINK
public static final String EXTERNAL_LINK
public static final String EXTERNAL_LINK_URL
public static final String CITATION
public static final String CATEGORY
public static final String BOLD
public static final String ITALICS
public static final String BOLD_ITALICS
public static final String HEADING
public static final String SUB_HEADING
public static final int ALPHANUM_ID
public static final int APOSTROPHE_ID
public static final int ACRONYM_ID
public static final int COMPANY_ID
public static final int EMAIL_ID
public static final int HOST_ID
public static final int NUM_ID
public static final int CJ_ID
public static final int INTERNAL_LINK_ID
public static final int EXTERNAL_LINK_ID
public static final int CITATION_ID
public static final int CATEGORY_ID
public static final int BOLD_ID
public static final int ITALICS_ID
public static final int BOLD_ITALICS_ID
public static final int HEADING_ID
public static final int SUB_HEADING_ID
public static final int EXTERNAL_LINK_URL_ID
public static final String[] TOKEN_TYPES
public static final String[] tokenImage
TOKEN_TYPES
insteadpublic static final int TOKENS_ONLY
public static final int UNTOKENIZED_ONLY
public static final int BOTH
public static final int UNTOKENIZED_TOKEN_FLAG
TOKENS_ONLY
was used, produce multiple tokens.
Constructor Detail |
---|
public WikipediaTokenizer(Reader input)
WikipediaTokenizer
. Attaches the
input
to a newly created JFlex scanner.
input
- The Input Readerpublic WikipediaTokenizer(Reader input, int tokenOutput, Set untokenizedTypes)
Method Detail |
---|
public Token next(Token reusableToken) throws IOException
TokenStream
This implicitly defines a "contract" between consumers (callers of this method) and producers (implementations of this method that are the source for tokens):
Token.clear()
before setting the fields in it & returning itTokenFilter
is considered a consumer.
next
in class TokenStream
reusableToken
- a Token that may or may not be used to
return; this parameter should never be null (the callee
is not required to check for null before using it, but it is a
good idea to assert that it is not null.)
IOException
public void reset() throws IOException
TokenStream
reset
in class TokenStream
IOException
public void reset(Reader reader) throws IOException
Tokenizer
reset
in class Tokenizer
IOException
|
||||||||||
PREV CLASS NEXT CLASS | FRAMES NO FRAMES | |||||||||
SUMMARY: NESTED | FIELD | CONSTR | METHOD | DETAIL: FIELD | CONSTR | METHOD |