Package com.itextpdf.io.source
Class PdfTokenizer
java.lang.Object
com.itextpdf.io.source.PdfTokenizer
- All Implemented Interfaces:
-
Closeable,AutoCloseable
Tokenizes PDF syntax from a random-access byte source.
Instances maintain a mutable stream position and token state and are not thread-safe.
-
Nested Class Summary
Nested ClassesModifier and TypeClassDescriptionstatic enumToken types recognized by this tokenizer. -
Field Summary
FieldsModifier and TypeFieldDescriptionstatic final byte[]static final byte[]protected intThe generation number parsed for the current indirect reference.protected booleanWhether the current string token uses hexadecimal notation.static final byte[]static final byte[]static final byte[]protected ByteBufferThe mutable bytes of the most recently parsed token.static final byte[]protected intThe object number parsed for the current indirect reference.static final byte[]static final byte[]static final byte[]static final byte[]protected PdfTokenizer.TokenTypeThe type of the most recently parsed token.static final byte[] -
Constructor Summary
Constructors -
Method Summary
Modifier and TypeMethodDescriptionvoidbackOnePosition(int ch) Pushes a read byte back so it becomes the next byte read.voidValidates that an FDF header begins at offset zero.static int[]checkObjectStart(PdfTokenizer lineTokenizer) Check whether line starts with object declaration.Validates a PDF header at offset zero and returns its version text.static booleancheckTrailer(ByteBuffer line) Checks whetherlineequals to 'trailer'.voidclose()Closes the underlying source when closing is enabled.static byte[]decodeStringContent(byte[] content, boolean hexWriting) Resolve escape symbols or hexadecimal symbols.protected static byte[]decodeStringContent(byte[] content, int from, int to, boolean hexWriting) Resolve escape symbols or hexadecimal symbols.byte[]Copies the bytes of the most recently parsed token.byte[]Decodes the current PDF string token.intgetGenNr()Gets the generation number parsed from the current indirect reference.intFinds a PDF or FDF header in the first kilobyte of the source.intParses the current token value as anint.longParses the current token value as along.longGets next %%EOF marker in current PDF file.intgetObjNr()Gets the object number parsed from the current indirect reference.longGets the current source position.Creates an independent view of the underlying source.longLocates the finalstartxrefmarker near the end of the source.Converts the current token bytes to a string using the platform default charset.Gets the type of the most recently parsed token.booleanTests whetherclose()closes the underlying source.protected static booleanisDelimiter(int ch) Tests whether a character is a PDF delimiter.protected static booleanisDelimiterWhitespace(int ch) Tests whether a character is a PDF delimiter or whitespace character.booleanTests whether the current string token uses hexadecimal notation.static booleanisWhitespace(int ch) Is a certain character a whitespace? Currently checks on the following: '0', '9', '10', '12', '13', '32'.protected static booleanisWhitespace(int ch, boolean isWhitespace) Checks whether a character is a whitespace.longlength()Gets the length of the underlying source.booleanParses the next PDF token into this tokenizer's mutable token state.voidReads the next non-comment token and recognizes indirect references and object declarations.intpeek()Gets the next byte of pdf source without moving source position.intpeek(byte[] buffer) Gets the nextbuffer.lengthbytes of pdf source without moving source position.intread()Reads one byte and advances the source position.voidreadFully(byte[] bytes) Reads enough bytes to fill a destination array.booleanreadLineSegment(ByteBuffer buffer) Reads data into the provided byte[].booleanreadLineSegment(ByteBuffer buffer, boolean isNullWhitespace) Reads data into the provided byte[].readString(int size) Reads up to a requested number of bytes as character values.voidseek(long pos) Sets the position from which the next byte is read.voidsetCloseStream(boolean closeStream) Configures whetherclose()closes the underlying source.voidthrowError(String error, Object... messageParams) Helper method to handle content errors.booleantokenValueEqualsTo(byte[] cmp) Tests whether the current token bytes equal a candidate byte sequence.
-
Field Details
-
Obj
public static final byte[] Obj -
R
public static final byte[] R -
Xref
public static final byte[] Xref -
Startxref
public static final byte[] Startxref -
Stream
public static final byte[] Stream -
Trailer
public static final byte[] Trailer -
N
public static final byte[] N -
F
public static final byte[] F -
Null
public static final byte[] Null -
True
public static final byte[] True -
False
public static final byte[] False -
type
The type of the most recently parsed token. -
reference
protected int referenceThe object number parsed for the current indirect reference. -
generation
protected int generationThe generation number parsed for the current indirect reference. -
hexString
protected boolean hexStringWhether the current string token uses hexadecimal notation. -
outBuf
The mutable bytes of the most recently parsed token.
-
-
Constructor Details
-
PdfTokenizer
Creates a PdfTokenizer for the specifiedRandomAccessFileOrArray. The beginning of the file is read to determine the location of the header, and the data source is adjusted as necessary to account for any junk that occurs in the byte source before the header- Parameters:
-
file- the source
-
-
Method Details
-
seek
public void seek(long pos) Sets the position from which the next byte is read.- Parameters:
-
pos- the absolute byte offset in the underlying source
-
readFully
Reads enough bytes to fill a destination array.- Parameters:
-
bytes- the destination array - Throws:
-
IOException- if EOF is reached or the source cannot be read
-
getPosition
public long getPosition()Gets the current source position.- Returns:
- the absolute offset of the next byte to read
-
close
Closes the underlying source when closing is enabled.- Specified by:
-
closein interfaceAutoCloseable - Specified by:
-
closein interfaceCloseable - Throws:
-
IOException- if the underlying source cannot be closed
-
length
public long length()Gets the length of the underlying source.- Returns:
- the number of readable bytes
-
read
Reads one byte and advances the source position.- Returns:
-
the unsigned byte value, or
-1at EOF - Throws:
-
IOException- if the source cannot be read
-
peek
Gets the next byte of pdf source without moving source position.- Returns:
- the byte, or -1 if EOF is reached
- Throws:
-
IOException- in case of any reading error
-
peek
Gets the nextbuffer.lengthbytes of pdf source without moving source position.- Parameters:
-
buffer- buffer to store read bytes - Returns:
-
the number of read bytes. If it is less than
buffer.lengthit means EOF has been reached - Throws:
-
IOException- in case of any reading error
-
readString
Reads up to a requested number of bytes as character values.- Parameters:
-
size- the maximum number of bytes to read - Returns:
- a string containing the bytes read before EOF
- Throws:
-
IOException- if the source cannot be read
-
getTokenType
Gets the type of the most recently parsed token.- Returns:
- the current token type
-
getByteContent
public byte[] getByteContent()Copies the bytes of the most recently parsed token.- Returns:
- a new array containing the current token bytes
-
getStringValue
Converts the current token bytes to a string using the platform default charset.- Returns:
- the current token value as a string
-
getDecodedStringContent
public byte[] getDecodedStringContent()Decodes the current PDF string token.- Returns:
- a new array containing decoded literal or hexadecimal string bytes
-
tokenValueEqualsTo
public boolean tokenValueEqualsTo(byte[] cmp) Tests whether the current token bytes equal a candidate byte sequence.- Parameters:
-
cmp- the bytes to compare;nullnever matches - Returns:
-
trueifcmpequals the current token bytes
-
getObjNr
public int getObjNr()Gets the object number parsed from the current indirect reference.- Returns:
- the parsed object number
-
getGenNr
public int getGenNr()Gets the generation number parsed from the current indirect reference.- Returns:
- the parsed generation number
-
backOnePosition
public void backOnePosition(int ch) Pushes a read byte back so it becomes the next byte read.- Parameters:
-
ch- the byte value to push back;-1is ignored
-
getHeaderOffset
Finds a PDF or FDF header in the first kilobyte of the source.- Returns:
- the zero-based byte offset of the header
- Throws:
-
IOException- if no supported header is found or the source cannot be read
-
checkPdfHeader
Validates a PDF header at offset zero and returns its version text.- Returns:
- the header text without its percent sign
- Throws:
-
IOException- if the header is absent or the source cannot be read
-
checkFdfHeader
Validates that an FDF header begins at offset zero.- Throws:
-
IOException- if the header is absent or the source cannot be read
-
getStartxref
Locates the finalstartxrefmarker near the end of the source.- Returns:
- the absolute byte offset of the marker
- Throws:
-
IOException- if the marker is absent or the source cannot be read
-
getNextEof
Gets next %%EOF marker in current PDF file.- Returns:
- next %%EOF marker position
- Throws:
-
IOException- in case of input-output related exceptions during PDF document reading
-
nextValidToken
Reads the next non-comment token and recognizes indirect references and object declarations.The source position and current token state are advanced to the recognized token.
- Throws:
-
IOException- if the source cannot be read or malformed syntax is encountered
-
nextToken
Parses the next PDF token into this tokenizer's mutable token state.- Returns:
-
truewhen a token was read, orfalseat EOF - Throws:
-
IOException- if the source cannot be read or token syntax is malformed
-
getLongValue
public long getLongValue()Parses the current token value as along.- Returns:
- the parsed numeric value
-
getIntValue
public int getIntValue()Parses the current token value as anint.- Returns:
- the parsed numeric value
-
isHexString
public boolean isHexString()Tests whether the current string token uses hexadecimal notation.- Returns:
-
truefor a hexadecimal string token
-
isCloseStream
public boolean isCloseStream()Tests whetherclose()closes the underlying source.- Returns:
-
trueif this tokenizer owns closing the source
-
setCloseStream
public void setCloseStream(boolean closeStream) Configures whetherclose()closes the underlying source.- Parameters:
-
closeStream-trueto close the source,falseto leave it open
-
getSafeFile
Creates an independent view of the underlying source.- Returns:
- a view with its own position; closing it does not close this tokenizer's source
-
decodeStringContent
protected static byte[] decodeStringContent(byte[] content, int from, int to, boolean hexWriting) Resolve escape symbols or hexadecimal symbols.NOTE Due to PdfReference 1.7 part 3.2.3 String value contain ASCII characters, so we can convert it directly to byte array.
- Parameters:
-
content- string bytes to be decoded -
from- given start index -
to- given end index -
hexWriting- true if given string is hex-encoded, e.g. '<69546578…>'. False otherwise, e.g. '((iText( some version)…)' - Returns:
-
byte[] for decrypting or for creating
String.
-
decodeStringContent
public static byte[] decodeStringContent(byte[] content, boolean hexWriting) Resolve escape symbols or hexadecimal symbols.
NOTE Due to PdfReference 1.7 part 3.2.3 String value contain ASCII characters, so we can convert it directly to byte array.- Parameters:
-
content- string bytes to be decoded -
hexWriting- true if given string is hex-encoded, e.g. '<69546578…>'. False otherwise, e.g. '((iText( some version)…)' - Returns:
-
byte[] for decrypting or for creating
String.
-
isWhitespace
public static boolean isWhitespace(int ch) Is a certain character a whitespace? Currently checks on the following: '0', '9', '10', '12', '13', '32'.
The same as callingisWhiteSpace(ch, true).- Parameters:
-
ch- int - Returns:
- boolean
-
isWhitespace
protected static boolean isWhitespace(int ch, boolean isWhitespace) Checks whether a character is a whitespace. Currently checks on the following: '0', '9', '10', '12', '13', '32'.- Parameters:
-
ch- int -
isWhitespace- boolean - Returns:
- boolean
-
isDelimiter
protected static boolean isDelimiter(int ch) Tests whether a character is a PDF delimiter.- Parameters:
-
ch- the character value to test - Returns:
-
truewhenchis a PDF delimiter
-
isDelimiterWhitespace
protected static boolean isDelimiterWhitespace(int ch) Tests whether a character is a PDF delimiter or whitespace character.- Parameters:
-
ch- the character value to test - Returns:
-
truewhenchis a delimiter or configured whitespace value
-
throwError
Helper method to handle content errors. Add file position toPdfRuntimeException.- Parameters:
-
error- message. -
messageParams- error params. - Throws:
-
IOException- wrap error message intoPdfRuntimeExceptionand add position in file.
-
checkTrailer
Checks whetherlineequals to 'trailer'.- Parameters:
-
line- for check - Returns:
- true, if line is equals to 'trailer', otherwise false
-
readLineSegment
Reads data into the provided byte[]. Checks on leading whitespace. SeeisWhiteSpace(int)orisWhiteSpace(int, boolean)for a list of whitespace characters.
The same as callingreadLineSegment(input, true).- Parameters:
-
buffer- aByteBufferto which the result of reading will be saved - Returns:
- true, if something was read or if the end of the input stream is not reached
- Throws:
-
IOException- in case of any reading error
-
readLineSegment
Reads data into the provided byte[]. Checks on leading whitespace. SeeisWhiteSpace(int)orisWhiteSpace(int, boolean)for a list of whitespace characters.- Parameters:
-
buffer- aByteBufferto which the result of reading will be saved -
isNullWhitespace- boolean to indicate whether '0' is whitespace or not. If in doubt, use true or overloaded methodreadLineSegment(input) - Returns:
- true, if something was read or if the end of the input stream is not reached
- Throws:
-
IOException- in case of any reading error
-
checkObjectStart
Check whether line starts with object declaration.- Parameters:
-
lineTokenizer- tokenizer, built by single line. - Returns:
- object number and generation if check is successful, otherwise - null.
-