View on GitHub

IRAF Community Distribution

IRAF maintained by the community

Home | Installation | Packages | X11IRAF | PyRAF | Forum

text files with long lines

Phil Hodge wrote on Mar 31, 1999

It appears that IRAF thinks that a file with lines longer than about 1024
characters is not a text file, even though it was opened as such.  Can
this limitation be removed?

I created a file using the following:

task	ttt

procedure ttt()

int	i
int	fd, open()

begin
	fd = open ("junk.txt", NEW_FILE, TEXT_FILE)

	do i = 1, 100 {
	    call fprintf (fd, "%15d")
		call pargi (i)
	}
	call fprintf (fd, "\n")

	call close (fd)
end

Then I tested the file type using dir l+ in the cl and also by calling
access, as in this test routine:

task	ttt

procedure ttt()

int	access()

begin
	if (access ("junk.txt", 0, TEXT_FILE) == YES)
	    call eprintf ("text file\n")
	else
	    call eprintf ("not a text file\n")
end

Doug Tody wrote on Mar 31, 1999

Hi Phil,

> It appears that IRAF thinks that a file with lines longer than about 1024
> characters is not a text file, even though it was opened as such.  Can
> this limitation be removed?
> 
> I created a file using the following:  [...]

The problem is deciding what is a "text file" and what is not.  This is
determined by the kernel routine zfacss.  Here are the comments in the
Unix version of the routine explaining what it is doing:

    /* If we have to check the file type (text or binary), then we must
     * actually look at some file data since UNIX does not discriminate
     * between text and binary files.  NOTE that this heuristic is not
     * completely reliable and can fail, although in practice it does
     * very well.
     */
    if (accessible && (acmode & R) && *type != 0) {
	stat ((char *)fname, &fi);

	/* Do NOT read from a special device (may block) */
	if ((fi.st_mode & S_IFMT) & S_IFREG) {
	    /* If we are testing for a text file the portion of the file
	     * tested must consist of only printable ascii characters or
	     * whitespace, with occasional newline line delimiters.
	     * Control characters embedded in the text will cause the
	     * heuristic to fail.  We require newlines to be present in
	     * the text to disinguish the case of a binary file containing
	     * only ascii data, e.g., a cardimage file.
	     */

For Unix (which does not distinguish text and binary files; we wouldn't
either if we only had to deal with Unix) the routine reads 1024 characters
of data from the start of the file for the test.  There has to be a newline
within 256 characters or the file is considered a binary file contain ascii
data, rather than a text file.

These numbers are arbitrary and can be adjusted.  In fact, since we 
increased the max line length in V2.11 to 1024 (1023+EOS), it appears
that zfacss ought to be updated as well, probably to increase everything
by 4 to match the rest of V2.11, e.g. permit a linelen of 1024, sampling
4096. 

Would that do it for you Phil?  To change this one routine in a V2.11
patch would be no problem.  To globally increase the line length for a
text file to more than 1024 would require a full recompile and extensive
testing, something that is not really feasible or safe for a patch.
IRAF can already handle data files with arbitrarily long lines, it is just
the standard text file stuff which is affected by this limit.

	- Doug

Phil Hodge wrote on Mar 31, 1999

Doug,

> The problem is deciding what is a "text file" and what is not.  This is
> determined by the kernel routine zfacss.  Here are the comments in the
> ...

> For Unix (which does not distinguish text and binary files; we wouldn't
> either if we only had to deal with Unix) the routine reads 1024 characters
> of data from the start of the file for the test.  There has to be a newline
> within 256 characters or the file is considered a binary file contain ascii
> data, rather than a text file.
>
> These numbers are arbitrary and can be adjusted.  In fact, since we 
> increased the max line length in V2.11 to 1024 (1023+EOS), it appears
> that zfacss ought to be updated as well, probably to increase everything
> by 4 to match the rest of V2.11, e.g. permit a linelen of 1024, sampling
> 4096. 

In the tables package, the current limit on line length for a text table
is 1024 (including newline), but occasionally someone will try to run an
IRAF task on an ascii file with longer lines than that, so I was thinking
of increasing the limit.  I'm currently using access() to distinguish text
tables from binary tables.  I do think it would be a good idea to increase
the length in zfacss, but for the purpose of checking table type I think the
solution is for me to do it differently, by examining the contents of the
file.  That is done eventually, of course, but it could be done initially
as well.  So I'll look into changing the table routines.

Phil

Last post on Mar 31, 1999