Class PdfTokenizer

java.lang.Object
com.itextpdf.io.source.PdfTokenizer
All Implemented Interfaces:
Closeable, AutoCloseable

public class PdfTokenizer extends Object implements Closeable
Tokenizes PDF syntax from a random-access byte source.

Instances maintain a mutable stream position and token state and are not thread-safe.

  • Nested Class Summary

    Nested Classes
    Modifier and Type
    Class
    Description
    static enum 
    Token types recognized by this tokenizer.
  • Field Summary

    Fields
    Modifier and Type
    Field
    Description
    static final byte[]
    F
     
    static final byte[]
     
    protected int
    The generation number parsed for the current indirect reference.
    protected boolean
    Whether the current string token uses hexadecimal notation.
    static final byte[]
    N
     
    static final byte[]
     
    static final byte[]
    Obj
     
    protected ByteBuffer
    The mutable bytes of the most recently parsed token.
    static final byte[]
    R
     
    protected int
    The object number parsed for the current indirect reference.
    static final byte[]
     
    static final byte[]
     
    static final byte[]
     
    static final byte[]
     
    The type of the most recently parsed token.
    static final byte[]
     
  • Constructor Summary

    Constructors
    Constructor
    Description
    Creates a PdfTokenizer for the specified RandomAccessFileOrArray.
  • Method Summary

    Modifier and Type
    Method
    Description
    void
    backOnePosition(int ch)
    Pushes a read byte back so it becomes the next byte read.
    void
    Validates that an FDF header begins at offset zero.
    static int[]
    checkObjectStart(PdfTokenizer lineTokenizer)
    Check whether line starts with object declaration.
    Validates a PDF header at offset zero and returns its version text.
    static boolean
    Checks whether line equals to 'trailer'.
    void
    close()
    Closes the underlying source when closing is enabled.
    static byte[]
    decodeStringContent(byte[] content, boolean hexWriting)
    Resolve escape symbols or hexadecimal symbols.
    protected static byte[]
    decodeStringContent(byte[] content, int from, int to, boolean hexWriting)
    Resolve escape symbols or hexadecimal symbols.
    byte[]
    Copies the bytes of the most recently parsed token.
    byte[]
    Decodes the current PDF string token.
    int
    Gets the generation number parsed from the current indirect reference.
    int
    Finds a PDF or FDF header in the first kilobyte of the source.
    int
    Parses the current token value as an int.
    long
    Parses the current token value as a long.
    long
    Gets next %%EOF marker in current PDF file.
    int
    Gets the object number parsed from the current indirect reference.
    long
    Gets the current source position.
    Creates an independent view of the underlying source.
    long
    Locates the final startxref marker near the end of the source.
    Converts the current token bytes to a string using the platform default charset.
    Gets the type of the most recently parsed token.
    boolean
    Tests whether close() closes the underlying source.
    protected static boolean
    isDelimiter(int ch)
    Tests whether a character is a PDF delimiter.
    protected static boolean
    Tests whether a character is a PDF delimiter or whitespace character.
    boolean
    Tests whether the current string token uses hexadecimal notation.
    static boolean
    isWhitespace(int ch)
    Is a certain character a whitespace? Currently checks on the following: '0', '9', '10', '12', '13', '32'.
    protected static boolean
    isWhitespace(int ch, boolean isWhitespace)
    Checks whether a character is a whitespace.
    long
    length()
    Gets the length of the underlying source.
    boolean
    Parses the next PDF token into this tokenizer's mutable token state.
    void
    Reads the next non-comment token and recognizes indirect references and object declarations.
    int
    peek()
    Gets the next byte of pdf source without moving source position.
    int
    peek(byte[] buffer)
    Gets the next buffer.length bytes of pdf source without moving source position.
    int
    read()
    Reads one byte and advances the source position.
    void
    readFully(byte[] bytes)
    Reads enough bytes to fill a destination array.
    boolean
    Reads data into the provided byte[].
    boolean
    readLineSegment(ByteBuffer buffer, boolean isNullWhitespace)
    Reads data into the provided byte[].
    readString(int size)
    Reads up to a requested number of bytes as character values.
    void
    seek(long pos)
    Sets the position from which the next byte is read.
    void
    setCloseStream(boolean closeStream)
    Configures whether close() closes the underlying source.
    void
    throwError(String error, Object... messageParams)
    Helper method to handle content errors.
    boolean
    tokenValueEqualsTo(byte[] cmp)
    Tests whether the current token bytes equal a candidate byte sequence.

    Methods inherited from class java.lang.Object

    clone, equals, finalize, getClass, hashCode, notify, notifyAll, toString, wait, wait, wait
  • Field Details

    • Obj

      public static final byte[] Obj
    • R

      public static final byte[] R
    • Xref

      public static final byte[] Xref
    • Startxref

      public static final byte[] Startxref
    • Stream

      public static final byte[] Stream
    • Trailer

      public static final byte[] Trailer
    • N

      public static final byte[] N
    • F

      public static final byte[] F
    • Null

      public static final byte[] Null
    • True

      public static final byte[] True
    • False

      public static final byte[] False
    • type

      protected PdfTokenizer.TokenType type
      The type of the most recently parsed token.
    • reference

      protected int reference
      The object number parsed for the current indirect reference.
    • generation

      protected int generation
      The generation number parsed for the current indirect reference.
    • hexString

      protected boolean hexString
      Whether the current string token uses hexadecimal notation.
    • outBuf

      protected ByteBuffer outBuf
      The mutable bytes of the most recently parsed token.
  • Constructor Details

    • PdfTokenizer

      public PdfTokenizer (RandomAccessFileOrArray file)
      Creates a PdfTokenizer for the specified RandomAccessFileOrArray. The beginning of the file is read to determine the location of the header, and the data source is adjusted as necessary to account for any junk that occurs in the byte source before the header
      Parameters:
      file - the source
  • Method Details

    • seek

      public void seek (long pos)
      Sets the position from which the next byte is read.
      Parameters:
      pos - the absolute byte offset in the underlying source
    • readFully

      public void readFully (byte[] bytes) throws IOException
      Reads enough bytes to fill a destination array.
      Parameters:
      bytes - the destination array
      Throws:
      IOException - if EOF is reached or the source cannot be read
    • getPosition

      public long getPosition()
      Gets the current source position.
      Returns:
      the absolute offset of the next byte to read
    • close

      public void close() throws IOException
      Closes the underlying source when closing is enabled.
      Specified by:
      close in interface AutoCloseable
      Specified by:
      close in interface Closeable
      Throws:
      IOException - if the underlying source cannot be closed
    • length

      public long length()
      Gets the length of the underlying source.
      Returns:
      the number of readable bytes
    • read

      public int read() throws IOException
      Reads one byte and advances the source position.
      Returns:
      the unsigned byte value, or -1 at EOF
      Throws:
      IOException - if the source cannot be read
    • peek

      public int peek() throws IOException
      Gets the next byte of pdf source without moving source position.
      Returns:
      the byte, or -1 if EOF is reached
      Throws:
      IOException - in case of any reading error
    • peek

      public int peek (byte[] buffer) throws IOException
      Gets the next buffer.length bytes of pdf source without moving source position.
      Parameters:
      buffer - buffer to store read bytes
      Returns:
      the number of read bytes. If it is less than buffer.length it means EOF has been reached
      Throws:
      IOException - in case of any reading error
    • readString

      public String readString (int size) throws IOException
      Reads up to a requested number of bytes as character values.
      Parameters:
      size - the maximum number of bytes to read
      Returns:
      a string containing the bytes read before EOF
      Throws:
      IOException - if the source cannot be read
    • getTokenType

      public PdfTokenizer.TokenType getTokenType()
      Gets the type of the most recently parsed token.
      Returns:
      the current token type
    • getByteContent

      public byte[] getByteContent()
      Copies the bytes of the most recently parsed token.
      Returns:
      a new array containing the current token bytes
    • getStringValue

      public String getStringValue()
      Converts the current token bytes to a string using the platform default charset.
      Returns:
      the current token value as a string
    • getDecodedStringContent

      public byte[] getDecodedStringContent()
      Decodes the current PDF string token.
      Returns:
      a new array containing decoded literal or hexadecimal string bytes
    • tokenValueEqualsTo

      public boolean tokenValueEqualsTo (byte[] cmp)
      Tests whether the current token bytes equal a candidate byte sequence.
      Parameters:
      cmp - the bytes to compare; null never matches
      Returns:
      true if cmp equals the current token bytes
    • getObjNr

      public int getObjNr()
      Gets the object number parsed from the current indirect reference.
      Returns:
      the parsed object number
    • getGenNr

      public int getGenNr()
      Gets the generation number parsed from the current indirect reference.
      Returns:
      the parsed generation number
    • backOnePosition

      public void backOnePosition (int ch)
      Pushes a read byte back so it becomes the next byte read.
      Parameters:
      ch - the byte value to push back; -1 is ignored
    • getHeaderOffset

      public int getHeaderOffset() throws IOException
      Finds a PDF or FDF header in the first kilobyte of the source.
      Returns:
      the zero-based byte offset of the header
      Throws:
      IOException - if no supported header is found or the source cannot be read
    • checkPdfHeader

      public String checkPdfHeader() throws IOException
      Validates a PDF header at offset zero and returns its version text.
      Returns:
      the header text without its percent sign
      Throws:
      IOException - if the header is absent or the source cannot be read
    • checkFdfHeader

      public void checkFdfHeader() throws IOException
      Validates that an FDF header begins at offset zero.
      Throws:
      IOException - if the header is absent or the source cannot be read
    • getStartxref

      public long getStartxref() throws IOException
      Locates the final startxref marker near the end of the source.
      Returns:
      the absolute byte offset of the marker
      Throws:
      IOException - if the marker is absent or the source cannot be read
    • getNextEof

      public long getNextEof() throws IOException
      Gets next %%EOF marker in current PDF file.
      Returns:
      next %%EOF marker position
      Throws:
      IOException - in case of input-output related exceptions during PDF document reading
    • nextValidToken

      public void nextValidToken() throws IOException
      Reads the next non-comment token and recognizes indirect references and object declarations.

      The source position and current token state are advanced to the recognized token.

      Throws:
      IOException - if the source cannot be read or malformed syntax is encountered
    • nextToken

      public boolean nextToken() throws IOException
      Parses the next PDF token into this tokenizer's mutable token state.
      Returns:
      true when a token was read, or false at EOF
      Throws:
      IOException - if the source cannot be read or token syntax is malformed
    • getLongValue

      public long getLongValue()
      Parses the current token value as a long.
      Returns:
      the parsed numeric value
    • getIntValue

      public int getIntValue()
      Parses the current token value as an int.
      Returns:
      the parsed numeric value
    • isHexString

      public boolean isHexString()
      Tests whether the current string token uses hexadecimal notation.
      Returns:
      true for a hexadecimal string token
    • isCloseStream

      public boolean isCloseStream()
      Tests whether close() closes the underlying source.
      Returns:
      true if this tokenizer owns closing the source
    • setCloseStream

      public void setCloseStream (boolean closeStream)
      Configures whether close() closes the underlying source.
      Parameters:
      closeStream - true to close the source, false to leave it open
    • getSafeFile

      public RandomAccessFileOrArray getSafeFile()
      Creates an independent view of the underlying source.
      Returns:
      a view with its own position; closing it does not close this tokenizer's source
    • decodeStringContent

      protected static byte[] decodeStringContent (byte[] content, int from, int to, boolean hexWriting)
      Resolve escape symbols or hexadecimal symbols.

      NOTE Due to PdfReference 1.7 part 3.2.3 String value contain ASCII characters, so we can convert it directly to byte array.

      Parameters:
      content - string bytes to be decoded
      from - given start index
      to - given end index
      hexWriting - true if given string is hex-encoded, e.g. '<69546578…>'. False otherwise, e.g. '((iText( some version)…)'
      Returns:
      byte[] for decrypting or for creating String.
    • decodeStringContent

      public static byte[] decodeStringContent (byte[] content, boolean hexWriting)
      Resolve escape symbols or hexadecimal symbols.
      NOTE Due to PdfReference 1.7 part 3.2.3 String value contain ASCII characters, so we can convert it directly to byte array.
      Parameters:
      content - string bytes to be decoded
      hexWriting - true if given string is hex-encoded, e.g. '<69546578…>'. False otherwise, e.g. '((iText( some version)…)'
      Returns:
      byte[] for decrypting or for creating String.
    • isWhitespace

      public static boolean isWhitespace (int ch)
      Is a certain character a whitespace? Currently checks on the following: '0', '9', '10', '12', '13', '32'.
      The same as calling isWhiteSpace(ch, true).
      Parameters:
      ch - int
      Returns:
      boolean
    • isWhitespace

      protected static boolean isWhitespace (int ch, boolean isWhitespace)
      Checks whether a character is a whitespace. Currently checks on the following: '0', '9', '10', '12', '13', '32'.
      Parameters:
      ch - int
      isWhitespace - boolean
      Returns:
      boolean
    • isDelimiter

      protected static boolean isDelimiter (int ch)
      Tests whether a character is a PDF delimiter.
      Parameters:
      ch - the character value to test
      Returns:
      true when ch is a PDF delimiter
    • isDelimiterWhitespace

      protected static boolean isDelimiterWhitespace (int ch)
      Tests whether a character is a PDF delimiter or whitespace character.
      Parameters:
      ch - the character value to test
      Returns:
      true when ch is a delimiter or configured whitespace value
    • throwError

      public void throwError (String error, Object... messageParams)
      Helper method to handle content errors. Add file position to PdfRuntimeException.
      Parameters:
      error - message.
      messageParams - error params.
      Throws:
      IOException - wrap error message into PdfRuntimeException and add position in file.
    • checkTrailer

      public static boolean checkTrailer (ByteBuffer line)
      Checks whether line equals to 'trailer'.
      Parameters:
      line - for check
      Returns:
      true, if line is equals to 'trailer', otherwise false
    • readLineSegment

      public boolean readLineSegment (ByteBuffer buffer) throws IOException
      Reads data into the provided byte[]. Checks on leading whitespace. See isWhiteSpace(int) or isWhiteSpace(int, boolean) for a list of whitespace characters.
      The same as calling readLineSegment(input, true).
      Parameters:
      buffer - a ByteBuffer to which the result of reading will be saved
      Returns:
      true, if something was read or if the end of the input stream is not reached
      Throws:
      IOException - in case of any reading error
    • readLineSegment

      public boolean readLineSegment (ByteBuffer buffer, boolean isNullWhitespace) throws IOException
      Reads data into the provided byte[]. Checks on leading whitespace. See isWhiteSpace(int) or isWhiteSpace(int, boolean) for a list of whitespace characters.
      Parameters:
      buffer - a ByteBuffer to which the result of reading will be saved
      isNullWhitespace - boolean to indicate whether '0' is whitespace or not. If in doubt, use true or overloaded method readLineSegment(input)
      Returns:
      true, if something was read or if the end of the input stream is not reached
      Throws:
      IOException - in case of any reading error
    • checkObjectStart

      public static int[] checkObjectStart (PdfTokenizer lineTokenizer)
      Check whether line starts with object declaration.
      Parameters:
      lineTokenizer - tokenizer, built by single line.
      Returns:
      object number and generation if check is successful, otherwise - null.